G1 Whole-Body Motion Tracking on Hardware¶
Hardware target
Unitree G1 humanoid (29-DoF variant). Joint order comes from the task owner’s
scene (src/unilab/assets/robots/g1/scene_flat.xml, actuator order); verify
that order against your SDK motor indices before hardware bring-up.
This guide covers the observation and action contract a G1 motion-tracking policy expects on hardware. The repository does not ship a G1 deploy runtime — you supply the hardware-side loop, and this page tells you what it must reproduce.
0. Verify your sim-side checkpoint¶
# Replay the policy headlessly and produce a video.
uv run eval --algo ppo --task g1_motion_tracking --sim motrix --load-run -1 \
--render-mode record
What to look for in the video:
The tracked bodies follow the reference motion without large discontinuities.
Joint velocities and actions remain finite and within the expected range.
Contact timing looks consistent with the reference motion.
If any of those is off, fix the sim-side checkpoint before hardware bring-up.
1. Export, then read the contract off the owner YAML¶
Export policy.onnx through the training playback path, using the same task
owner that produced the checkpoint:
uv run eval --algo ppo --task g1_motion_tracking --sim motrix --load-run -1
Every field your hardware loop needs is declared in that owner’s YAML. Widths differ per owner — two G1 examples:
Owner |
Actor obs width |
Notes |
|---|---|---|
|
514 |
No state estimation: |
|
154 |
Single-step mimic actor layout, per-joint-group |
Read the width off the composed config, not off this table
Actor obs width is the sum of dim * history_length over the terms declared
under env.observations.actor.terms, plus the motion command. A term set to
null is dropped. If the ONNX input width disagrees with what your hardware
loop assembles, that is a contract bug, not a hardware tuning problem.
2. Observation contract¶
Term order and per-term history come from your own owner’s
env.observations.actor.terms. As a worked example, g1_wbt_obs declares the
terms below; those carrying history_length: 5 are flattened oldest-first
within the term, and terms are concatenated in declaration order:
Term |
Dim |
Source on hardware |
|---|---|---|
|
6 |
anchor orientation term from the reference and robot torso frames |
|
3 per history step |
IMU gyro ( |
|
29 per history step |
measured joint position minus the |
|
29 per history step |
joint velocity term |
|
29 per history step |
previous raw actor output |
The motion command contributes the reference joint position and velocity
(29 + 29) ahead of the observation terms. Per-term oldest-first ordering is
guarded by tests/scripts/test_obs_alignment_g1_wbt.py; mirror that ordering
on hardware or the policy reads a permuted vector.
3. Actuator interface¶
Map actor output as action * scale + default_angles, then clamp to the
scene’s joint range before the target reaches the motor driver.
scaleisenv.actions.joint_pos.scale. It may be a scalar (2.0forg1_wbt_obs) or a regex → value map resolved per actuator (g1_motion_tracking_deploymaps joint-name patterns to distinct values). Reproduce the owner’s resolved per-actuator vector exactly — do not average a map, take one entry, or broadcast a scalar over a map owner.default_anglesfollows fromuse_default_offset: true, i.e. thestandkeyframe joint block of the owner’s scene.Joint limits and gains come from the same scene XML (
jnt_range, position actuatorgainprm/biasprm).
Training applies the target directly with no smoothing. If hardware jitter forces you to add smoothing, verify the sim2sim impact first — every step of lag pushes observations out of the training distribution.
4. Reference motion sync¶
The phase variable lets the policy track an externally-supplied motion clip. On hardware you need a wall-clock → phase mapping that is:
Monotonic — no skipping back.
Restartable — survives a comms hiccup without producing a step discontinuity in
(sin φ, cos φ).Bounded rate — clip dφ/dt to the value the policy was trained with (the motion loader records this; load
reference_motion.npz).
See unilab.tasks.motion_tracking.common.motion_loader for the sim-side
loader you should mirror on hardware.
5. Safety layer¶
Hardware-side: see Hardware Safety Layers for the standard structure. The G1 specifics:
Reject non-finite actions and shape mismatches before applying the action scale.
Clamp generated targets with the joint range from the owner’s scene XML.
Keep watchdog, pose monitor, and operator-stop thresholds in the deploy controller and test them independently of the policy.
6. Closed-loop bring-up sequence¶
Stand-on-stand. Robot held by a gantry. Policy runs but actuators are torque-disabled. Confirm observation pipeline.
Torque-enable, hand-held. Operator catches the robot. Policy commands actuators. Confirm action mapping.
Gantry-supported gait. Track motion at half time-rate (dφ/dt halved).
Free-stand. Full rate, then remove gantry.
Do not skip the observation-only stage: it is where axis-order, joint-order,
and actions wiring mistakes are easiest to catch.
7. What to log¶
Log the full observation vector, full action vector, and wall clock for every step. Compare the first hardware observation window against a sim episode built from the same owner YAML — that diff localizes unit, frame, and ordering mistakes faster than any reward inspection.