Work / Humanoid robotics Simulation

Humanoid control

Unitree G1 whole-body control with GR00T and SONIC

Simulated G1 performing a pen-drawing task.Simulation

This work builds a loco-manipulation stack for the Unitree G1 humanoid. NVIDIA GR00T acts as the vision-language-action brain and outputs latent motion tokens, while GEAR-SONIC acts as the whole-body controller that turns those tokens into coordinated leg, waist, arm and hand motion.

Challenge

Humanoid loco-manipulation needs a policy that understands tasks and a controller that keeps the robot balanced while the arms do the work. Neither can be tuned safely on hardware first.

Approach

Development follows a simulation-first path: bring up the SONIC sim-to-sim loop, generate demonstrations in simulation, fine-tune the policy and evaluate before moving to hardware. The work included research into motion and retargeting datasets (BONES-SEED, GRAIL, OmniRetarget) and the Unitree SDK ecosystem. Alongside it, motion-capture retargeting drives upper-body motion on the simulated G1.

In brief

  • GR00T VLA with the SONIC whole-body controller on the Unitree G1
  • Simulation-first pipeline in Isaac Lab and MuJoCo
  • Synthetic egocentric manipulation dataset of 82 validated episodes
  • Dataset research: BONES-SEED, GRAIL, OmniRetarget
  • Motion-capture retargeting for upper-body motion

Egocentric manipulation dataset

To train the vision-language-action policy, the lab built a synthetic egocentric manipulation dataset: 82 validated episodes recorded from the G1's own head camera while the whole-body controller keeps the robot balanced. Tasks include picking fruit and placing it on a plate or tray, stacking and dropping onto a block, left-hand variants, lifting a board with both hands, pouring from a bottle or mug into a bowl, and opening a drawer. Each episode carries its language instruction, so the same data trains instruction following and manipulation together.

Eight episodes from the dataset, each captioned with its task instruction.Simulation

Locomotion and whole-body control

The whole-body controller handles balance, walking and turning while the policy commands the task; the same stack walks the robot to a goal before manipulation starts.