ICRA 2027 · Supplementary material

Whole-Body Dexterity

Anonymous Authors

G1 humanoid repositioning a basketball and box across its body, and stowing a box under one arm during a shelf task
Whole-Body Dexterity (WBD) extends manipulation beyond the hands by using body links as contact surfaces to hold objects. Changing these contacts allows objects to be repositioned across the body and enables natural human manipulation such as stowing objects underarm to free up other limbs.

Abstract

Humanoid loco-manipulation remains predominantly hand-centric, falling short of humans’ ability to engage the entire body by bracing and sliding bulky objects against the forearms and torso. We introduce Whole-Body Dexterity (WBD), a general framework that uses the robot's body and appendages as contact surfaces to grasp and manipulate objects. Our sim-to-real pipeline first generates a cache of feasible whole-body grasps, then trains reinforcement-learning experts on a general any-to-any grasp objective by sampling initial and goal grasps from the cache, and finally distills these experts into a single deployable policy. We show that the distilled policy transfers zero-shot to a real Unitree G1, where it repositions a basketball to regions around the torso and completes a long-horizon task that stows a box under one arm, carries it between shelves, and places it. The controller operates interactively, steered by coarse motion references from a human teleoperator. In simulation, a similar controller refines corrupted expert commands into stable manipulation. Across real-world whole-body manipulation tasks, the distilled controller completes 6× as many trials as direct retargeting, reaches target regions about 3× faster among successful trials, and completes tasks a generalist whole-body controller cannot.

Method

Pipeline: mode-based grasp dataset, sphere expert pre-training, per-object post-training, and distillation to a general dexterity controller deployed in sim and real
Our three-stage pipeline. (1) Generate a large corpus of whole-body grasps with energy-based optimization, subdividing the torso into contact regions. (2) Train per-object experts on any-to-any regrasping, warm-started from the grasp corpus, encouraging sliding and rolling contact across the torso. (3) Distill the experts into a deployable controller without privileged information.

Whole-Body Grasp Generation

Whole-body grasps across arms, legs and torso generated by energy-based optimization
Whole-Body Grasp Generation. Energy-based optimization generates grasp candidates from specified arm, leg, and torso contact links (blue), followed by simulation validation under perturbations. The voxelized training set contains left-arm, right-arm, and both-arm grasps; leg grasps are excluded.

Optimization

Simulation filter

Arm–torso

Leg–arm

Grasps are optimized by minimizing a differentiable force-closure residual over the object's signed distance field from the specified contact links (blue). Each candidate is then pressed into contact and held under 4× gravity while a sustained wrench disturbance is swept over the six axis directions; it is kept only if contact persists and the object then settles under standard gravity. Leg–arm grasps are generated to show the generality of the procedure but are not used for WBD-EX training.

WBD-EX Training

Voxelized grasp sampling

Initial and goal grasps are sampled from the voxelized grasp set around the torso, and the expert repositions the object from one grasp to the other.

Sphere pre-training → object post-training

A sphere expert is pre-trained on the any-to-any objective, then post-trained into experts for other object primitives.

Simulation Results

Noisy expert

Noisy expert + WBD

A paired trial from the same start. The translucent red robot is the noisy expert command and the blue cube is the goal. Executed directly, the noisy command drops the object; the WBD controller refines the same command into stable manipulation.

Goals per trial and failure-free time versus guidance for the noisy expert with and without WBD refinement, at two noise levels
Simulated object repositioning under noisy commands. Dashed lines show the noisy expert; solid curves show the noisy expert with WBD refinement. Training objects are weighted equally. Bands and error bars denote 95% paired bootstrap confidence intervals over 256 matched grasp starts per object.

Real-World Results

Real-world shelf task, shoulder tap and armpit stow with a basketball, and the Xsens teleoperation suit
Real-world whole-body dexterity. On-body repositioning and a shelf pick, stow, carry, un-stow, and place sequence using the WBD controller.

Experimental setup: teleoperation with WBD controller

Teleoperator (Xsens suit)

Teleop + WBD controller

The distilled controller is steered interactively by coarse motion references from a human teleoperator.

On-body repositioning (2× speed)

SONIC

GMR

WBD (ours)

Shoulder touch

Armpit stow

TaskControllerRegion reached ↑Full round trip ↑Time to region (s) ↓Total time (s) ↓
Shoulder touchSONIC2/100/1015.2—
GMR3/101/1026.030.1
WBD10/109/107.717.3
Armpit stowSONIC1/100/1028.0—
GMR4/102/1031.549.8
WBD10/109/1010.626.2

Ten trials per controller and task: region-reaching and full round-trip success. Times are measured from trial start and averaged over successful trials.

Long-horizon task: shelf pick, stow, carry and place

SONIC

GMR

WBD (ours)

Playback speed is shown in the top-right corner of each clip.

ControllerPick ↑Stow ↑Carry ↑Un-stow ↑Place ↑
SONIC10/100/100/100/100/10
GMR10/103/102/101/101/10
WBD10/109/109/109/107/10

Ten trials per controller: completed stages of the shelf task. Place denotes full-task success.