Skip to content
humanoidrobots.training

Booster T1 · runs in the tab

Velocity tracking with a gait clock (forward, lateral, yaw rate)

Record

Robot
Booster T1 booster/t1
Framework
booster_gym (Isaac Gym) PPO, MLP 47-256-128-128-12, ELU · Isaac Gym
Policy license
Apache-2.0
File
deploy/models/T1.pt
Contract
47 obs → 12 actions @ 50 Hz

Our walk test · 2026-09-25

Runtime
MuJoCo 3.14.0 WASM + onnxruntime-web 1.30 (wasm, 1 thread), Chromium, real time
Command
vx 0.5 m/s · vy 0 m/s · wz 0 rad/s
Duration
60 s
Mean speed
0.438 m/s
Min base z
0.642 m
Falls
0
Export
TorchScript of model.actor (booster_gym export_model.py); play_mujoco.py uses dist.loc = actor(obs), the same mean action → ONNX opset 17, weights unchangedparity: max |a_torch - a_onnx| < 3e-6 over 200 random inputs (onnxruntime 1.x CPU)
On hardware
booster_gym's real-robot deploy config (deploy/configs/T1.yaml) loads this same models/T1.pt, and the README links a video of the deploy on a physical T1. authors' guide
Files
policy.onnx · contract.jsononnx sha256 14f3e1ca7e3caf8b7b3d9eb22e373b9fc389b182559ff01914473625632754ec

Observation / action contract

47 obs → 12 actions @ 50 Hz · dt 0.002 s × 10 · infer before-step · action scale 1 · clip ±1

Copied from play_mujoco.py at da396a06d6 with envs/T1.yaml. ONNX input obs, output actions.

Observation terms in order
IndexTermSizeScale / params
0–2projected_gravity · sensor orientation31
3–5base_ang_vel · sensor angular-velocity31
6–8command31, 1, 1
9–10gait_clock21.5 Hz
11–22joint_pos_rel121
23–34joint_vel120.1
35–46last_action12·
Actuated joints in policy order with PD gains and default angles
#JointDefault radKpKd
0Left_Hip_Pitch-0.2002005
1Left_Hip_Roll0.0002005
2Left_Hip_Yaw0.0002005
3Left_Knee_Pitch0.4002005
4Left_Ankle_Pitch-0.250501
5Left_Ankle_Roll0.000501
6Right_Hip_Pitch-0.2002005
7Right_Hip_Roll0.0002005
8Right_Hip_Yaw0.0002005
9Right_Knee_Pitch0.4002005
10Right_Ankle_Pitch-0.250501
11Right_Ankle_Roll0.000501
vx trained / offered
[-1, 1] / [-0.6, 1]
vy trained / offered
[-1, 1] / [-0.5, 0.5]
wz trained / offered
[-1, 1] / [-0.8, 0.8]
  • Loop order follows play_mujoco.py: at every 10th step (from step 0) the policy runs on the current state, then PD, then mj_step, then the gait clock advances by dt * f.
  • Gravity and angular velocity come from the MJCF's framequat 'orientation' and gyro 'angular-velocity' sensors on the trunk IMU site, read from sensordata exactly as the script does.
  • PD torque is clipped to each motor's ctrlrange (MuJoCo clamps limited ctrl). Actions are clipped to +-1 before use and fed back as the last-action term.
  • Gait frequency is the mean of the trained range [1.0, 2.0] Hz while any command is non-zero, and 0 (clock terms zeroed) when all commands are zero, as in play_mujoco.py.
  • The sensors declare noise in the MJCF; MuJoCo does not apply sensor noise, so the replay is deterministic.
  • Replay check: this runtime (MuJoCo WASM + onnxruntime-web in Node) and the reference script (MuJoCo 3.14 Python + TorchScript) give the same base trajectory to the centimetre over 30 s at [0.5, 0, 0] (x 10.09 m, y -7.19 m). At that command the policy curves to the right in the reference script as well; steer with yaw.