Result videos · simulation (MuJoCo) · anonymous submission
robot driven by the shared policyreference motion (retargeted mocap)
One shared network controls every robot shown here. It is conditioned on joint and body descriptors, the current joint state and the motion input; the network has no robot-specific layers. Videos autoplay and loop; the camera follows the robot.
The shared policy (left) next to the reference it follows (right), on the training motions.
Unitree H1
Unitree G1
Booster T1
Boston Dynamics Atlas
Apptronik Apollo
A LAFAN1 dance sequence that was never used for training, on the same five-robot policy.
Unitree H1, held-out dance
Unitree G1, held-out dance
Booster T1, held-out dance
Atlas, held-out dance
Apollo, held-out dance
Unitree H1, Zeibekiko (DanceDB motion capture)
Unitree G1, Bachata (DanceDB motion capture)
Every training episode samples a new body variant (link lengths, masses, joint axes, actuator gains). Here the five-robot policy runs on randomized bodies at the strongest randomization level used in training; meshes keep their nominal size, so only the changed link offsets are visible.
Unitree H1: three randomized bodies, one policy, each next to its own reference (drawn on the nominal body).
Unitree G1: three randomized bodies, one policy, each next to its own reference.
The same on Atlas: three randomized bodies, one policy, each next to its own reference.
Two five-robot policies trained identically except for the motion input.
Left: the policy that sees only the current target pose loses the arm motion after long training under body randomization. Middle: adding a short window of the motion (as a token) keeps it locked to the reference.
The same two five-robot policies on the held-out dance: reference-only (left) vs. token (middle).
The same effect on a two-robot policy trained four times longer: reference-only (left) vs. token (middle).
The policy receives no separate current joint-target input. Its motion command is an FSQ code requiring 96 bits per joint and update if bit-packed.
Unitree G1
Booster T1
Boston Dynamics Atlas
Other segments of the 20 training motions (dances, walks, runs, jumps, fights).
Unitree H1
Unitree G1
Booster T1
Boston Dynamics Atlas
A four-robot policy (H1, G1, T1, Atlas) applied to Apollo without further training. It still receives Apollo's embodiment descriptors and retargeted reference; this is a held-out-body test, not code-only transfer.
The policy starts to follow the motion but falls within about a second. Transfer to an unseen robot is preliminary.