Abstract
Humanoid loco-manipulation remains predominantly hand-centric, falling short of humans’ ability to engage the entire body by bracing and sliding bulky objects against the forearms and torso. We introduce Whole-Body Dexterity (WBD), a general framework that uses the robot's body and appendages as contact surfaces to grasp and manipulate objects. Our sim-to-real pipeline first generates a cache of feasible whole-body grasps, then trains reinforcement-learning experts on a general any-to-any grasp objective by sampling initial and goal grasps from the cache, and finally distills these experts into a single deployable policy. We show that the distilled policy transfers zero-shot to a real Unitree G1, where it repositions a basketball to regions around the torso and completes a long-horizon task that stows a box under one arm, carries it between shelves, and places it. The controller operates interactively, steered by coarse motion references from a human teleoperator. In simulation, a similar controller refines corrupted expert commands into stable manipulation. Across real-world whole-body manipulation tasks, the distilled controller completes 6× as many trials as direct retargeting, reaches target regions about 3× faster among successful trials, and completes tasks a generalist whole-body controller cannot.
Method
Whole-Body Grasp Generation
Optimization
Simulation filter
Arm–torso
Leg–arm
Grasps are optimized by minimizing a differentiable force-closure residual over the object's signed distance field from the specified contact links (blue). Each candidate is then pressed into contact and held under 4× gravity while a sustained wrench disturbance is swept over the six axis directions; it is kept only if contact persists and the object then settles under standard gravity. Leg–arm grasps are generated to show the generality of the procedure but are not used for WBD-EX training.
WBD-EX Training
Voxelized grasp sampling
Sphere pre-training → object post-training
Simulation Results
Noisy expert
Noisy expert + WBD
A paired trial from the same start. The translucent red robot is the noisy expert command and the blue cube is the goal. Executed directly, the noisy command drops the object; the WBD controller refines the same command into stable manipulation.
Real-World Results
Experimental setup: teleoperation with WBD controller
Teleoperator (Xsens suit)
Teleop + WBD controller
The distilled controller is steered interactively by coarse motion references from a human teleoperator.
On-body repositioning (2× speed)
SONIC
GMR
WBD (ours)
Shoulder touch
Armpit stow
| Task | Controller | Region reached ↑ | Full round trip ↑ | Time to region (s) ↓ | Total time (s) ↓ |
|---|---|---|---|---|---|
| Shoulder touch | SONIC | 2/10 | 0/10 | 15.2 | — |
| GMR | 3/10 | 1/10 | 26.0 | 30.1 | |
| WBD | 10/10 | 9/10 | 7.7 | 17.3 | |
| Armpit stow | SONIC | 1/10 | 0/10 | 28.0 | — |
| GMR | 4/10 | 2/10 | 31.5 | 49.8 | |
| WBD | 10/10 | 9/10 | 10.6 | 26.2 |
Ten trials per controller and task: region-reaching and full round-trip success. Times are measured from trial start and averaged over successful trials.
Long-horizon task: shelf pick, stow, carry and place
SONIC
GMR
WBD (ours)
Playback speed is shown in the top-right corner of each clip.
| Controller | Pick ↑ | Stow ↑ | Carry ↑ | Un-stow ↑ | Place ↑ |
|---|---|---|---|---|---|
| SONIC | 10/10 | 0/10 | 0/10 | 0/10 | 0/10 |
| GMR | 10/10 | 3/10 | 2/10 | 1/10 | 1/10 |
| WBD | 10/10 | 9/10 | 9/10 | 9/10 | 7/10 |
Ten trials per controller: completed stages of the shelf task. Place denotes full-task success.