Skip to content
humanoidrobots.training

Unitree G1 · runs in the tab

Velocity tracking on flat ground (forward, lateral, yaw rate)

Record

Robot
Unitree G1 unitree/g1
Framework
legged_gym (Isaac Gym) + rsl_rl PPO, ActorCriticRecurrent (LSTM 64, MLP 32, ELU) · Isaac Gym
Policy license
BSD-3-Clause
File
deploy/pre_train/g1/motion.pt
Contract
47 obs → 12 actions @ 50 Hz · LSTM

Our walk test · 2026-09-25

Runtime
MuJoCo 3.14.0 WASM + onnxruntime-web 1.30 (wasm, 1 thread), Chromium, real time
Command
vx 0.5 m/s · vy 0 m/s · wz 0 rad/s
Duration
60 s
Mean speed
0.474 m/s
Min base z
0.763 m
Falls
0
Export
TorchScript PolicyExporterLSTM (hidden/cell state kept as module buffers) → ONNX opset 17 with the LSTM state as explicit inputs/outputs; weights copied unchangedparity: max |a_torch - a_onnx| < 5e-6 over 200 stateful random steps (onnxruntime 1.x CPU)
On hardware
unitree_rl_gym's deploy_real config for Unitree G1 loads this same checkpoint (g1.yaml policy_path) and the deploy guide shows it on the physical robot. authors' guide
Files
policy.onnx · contract.jsononnx sha256 64451da5909cc03a96fb8e435dc2495217292614466f7628718d60a6cab834fd

Observation / action contract

47 obs → 12 actions @ 50 Hz · LSTM · dt 0.002 s × 10 · infer after-step · action scale 0.25

Copied from deploy/deploy_mujoco/deploy_mujoco.py at 276801e46c with deploy/deploy_mujoco/configs/g1.yaml. ONNX input obs, output actions, recurrent state h_in/c_in [1×1×64].

Observation terms in order
IndexTermSizeScale / params
0–2base_ang_vel30.25
3–5projected_gravity3·
6–8command32, 2, 0.25
9–20joint_pos_rel121
21–32joint_vel120.05
33–44last_action12·
45–46gait_phase2period 0.8 s
Actuated joints in policy order with PD gains and default angles
#JointDefault radKpKd
0left_hip_pitch_joint-0.1001002
1left_hip_roll_joint0.0001002
2left_hip_yaw_joint0.0001002
3left_knee_joint0.3001504
4left_ankle_pitch_joint-0.200402
5left_ankle_roll_joint0.000402
6right_hip_pitch_joint-0.1001002
7right_hip_roll_joint0.0001002
8right_hip_yaw_joint0.0001002
9right_knee_joint0.3001504
10right_ankle_pitch_joint-0.200402
11right_ankle_roll_joint0.000402
vx trained / offered
[-1, 1] / [-0.6, 1]
vy trained / offered
[-1, 1] / [-0.5, 0.5]
wz trained / offered
[-1, 1] / [-0.8, 0.8]
  • The reference script runs PD on every 2 ms physics step: tau = kp*(target - q) - kd*qdot, written straight to the motor ctrl; MuJoCo clamps it to each joint's actuatorfrcrange.
  • The policy runs after every 10th step on the post-step state (50 Hz). gait phase = (steps*dt mod 0.8)/0.8, steps counted from reset.
  • Base angular velocity is qvel[3:6] of the free joint (body frame). Projected gravity uses the reference formula on qpos[3:7] (w,x,y,z).
  • Training used heading mode (yaw-rate command recomputed from heading error); the policy input is still (vx, vy, yaw rate).
  • Replay check: this runtime (MuJoCo WASM + onnxruntime-web in Node) and the reference script (MuJoCo 3.14 Python + TorchScript) give the same base trajectory to the centimetre over 30 s at [0.5, 0, 0] (x 13.69 m, y -3.03 m); the policy drifts slightly right there too.