Motion-Window Conditioning for Embodiment-Aware Multi-Humanoid Motion Tracking

Result videos · simulation (MuJoCo) · anonymous submission

Unitree H1
Unitree G1
Booster T1
Atlas
Apollo

robot driven by the shared policyreference motion (retargeted mocap)

One network controls every robot shown here. It receives a description of the robot's joints and body, the joint state and a short window of the target motion; nothing in it is specific to a robot. Videos autoplay and loop; the camera follows the robot.

One policy, five humanoids

The shared policy (left) next to the reference it follows (right), on the training motions.

Unitree H1

Unitree G1

Booster T1

Boston Dynamics Atlas

Apptronik Apollo

Held-out dances

A LAFAN1 dance sequence that was never used for training, on the same five-robot policy.

Unitree H1, held-out dance

Unitree G1, held-out dance

Booster T1, held-out dance

Atlas, held-out dance

Apollo, held-out dance

Unitree H1, Zeibekiko (DanceDB motion capture)

Unitree G1, Bachata (DanceDB motion capture)

Trained with body randomization

Every training episode samples a new body variant (link lengths, masses, joint axes, actuator gains). Here the five-robot policy runs on randomized bodies at the strongest randomization level used in training; meshes keep their nominal size, so only the changed link offsets are visible.

Unitree H1: three randomized bodies, one policy, each next to its own reference (drawn on the nominal body).

Unitree G1: three randomized bodies, one policy, each next to its own reference.

The same on Atlas: three randomized bodies, one policy, each next to its own reference.

Why the motion input matters

Two five-robot policies trained identically except for the motion input.

Left: the policy that sees only the current target pose loses the arm motion after long training under body randomization. Middle: adding a short window of the motion (as a token) keeps it locked to the reference.

The same two five-robot policies on the held-out dance: reference-only (left) vs. token (middle).

The same effect on a two-robot policy trained four times longer: reference-only (left) vs. token (middle).

Driven by codes alone

The policy never sees a joint target; its only motion input is a 96-bit code per joint and frame.

Unitree G1

Booster T1

Boston Dynamics Atlas

More training motions

Other segments of the 20 training motions (dances, walks, runs, jumps, fights).

Unitree H1

Unitree G1

Booster T1

Boston Dynamics Atlas

A robot never seen in training

A four-robot policy (H1, G1, T1, Atlas) applied to Apollo without any training on it; the only input it can transfer through is the per-joint code.

The policy starts to follow the motion but falls within about a second. Transfer to an unseen robot is preliminary.