bench.humanoidrobots.training · stand-hold-v1
Does the model stand?
Before you train a policy on a robot model, the model has to load and hold still. Every published MJCF in the browser sim is compiled headless in MuJoCo 3.14.0 (the same WebAssembly build the tab runs) and held for 10 s under one generic controller. A fall here is a fact about the model plus our controller, not a verdict on the robot.
Models
15
Hold a stand
11/15
Fell
4
Run
2026-09-24
Runner
32 cores · local
Results are 13 days old; the nightly run has not landed since.
Results
Sort by any column · default: passing first · RTF depends on the machine · results.json
| Apptronik Apolloapptronik_apollo | pass | 98.5% | 98.8% | -1.1 cm | 1.5° | 55.5× | 2.98 s | 80.9 kg | 32 |
| Berkeley Humanoidberkeley_humanoid | pass | 97.9% | 97.9% | -1.2 cm | 4.5° | 49.0× | 1.30 s | 16.1 kg | 12 |
| Booster T1booster_t1 | pass | 99.1% | 99.5% | -0.0 cm | 4.9° | 29.5× | 0.41 s | 31.6 kg | 23 |
| Booster T1 12-DoFbooster_t1_gym | pass | 94.8% | 95.6% | -2.7 cm | 0.3° | 51.8× | 0.39 s | 31.6 kg | 12 |
| Fourier N1fourier_n1 | pass | 99.8% | 99.8% | -0.1 cm | 0.7° | 38.2× | 2.69 s | 39.7 kg | 23 |
| PAL TALOSpal_talos | pass | 98.8% | 99.4% | -0.4 cm | 1.9° | 6.0× | 0.72 s | 94.0 kg | 32 |
| ROBOTIS OP3robotis_op3 | pass | 92.5% | 92.9% | -2.1 cm | 0.6° | 33.6× | 1.87 s | 3.1 kg | 20 |
| Unitree G1unitree_g1 | pass | 100.0% | 100.2% | 0.2 cm | 0.0° | 26.4× | 1.09 s | 33.3 kg | 29 |
| Unitree G1 12-DoFunitree_g1_12dof | pass | 99.6% | 99.7% | -0.2 cm | 0.8° | 42.2× | 1.14 s | 32.1 kg | 12 |
| Unitree H1unitree_h1 | pass | 99.6% | 99.6% | -0.4 cm | 0.7° | 41.9× | 0.56 s | 51.4 kg | 19 |
| Unitree H1 10-DoFunitree_h1_10dof | pass | 98.8% | 98.8% | -1.4 cm | 0.8° | 56.0× | 0.71 s | 51.6 kg | 10 |
| Agility Cassieagility_cassie | fell | 17.0% | 17.0% | -84.8 cm | 88.3° | 8.9× | 0.46 s | 33.3 kg | 10 |
| Dropbeardropbear · meshes simplified | fell | 13.3% | 13.3% | 23.3 cm | 103.0° | 0.6× | 8.14 s | 61.3 kg | 19 |
| PNDbotics Adam Litepndbotics_adam_lite | fell | 11.2% | 11.4% | -83.0 cm | 90.1° | 55.1× | 4.13 s | 58.2 kg | 25 |
| ToddlerBot 2XCtoddlerbot_2xc | fell | 32.4% | 32.5% | -19.8 cm | 149.8° | 15.9× | 2.52 s | 3.5 kg | 30 |
Method · stand-hold-v1
- Controller
- The browser sim's hold controller (packages/sim/src/engine.ts), mirrored headless: joint PD toward the stand keyframe (or qpos0) plus an ankle-strategy balance term from torso tilt, per-model gains from the @hs/sim manifest. No policy, no harness, no pushes.
- Pass
- Runs the full 10 s without a physics reset, the centre of mass never drops below 80% of its start height, and it ends within 10% of the start. Base drift is reported alongside.
- Fell
- The centre of mass drops below 60% of its start height at any point, or the physics diverges (NaN / auto-reset). Heights use the whole-body centre of mass because some models put the base origin near the floor.
- Real-time factor
- Simulated seconds divided by wall seconds for the 10 s run, single thread, no rendering. Depends on the machine; see environment.
- Load time
- Compile time: writing the pinned files into the WASM filesystem, mesh simplification when the model's runtime says so, MjModel.from_xml_path, and controller setup. Download time is excluded.
- Models
- Pinned to a commit of the model's own repo (see each model page). Some large meshes are simplified to fit the WASM heap, flagged per model.
Environment
- MuJoCo
- 3.14.0 (WASM)
- Node
- v23.11.0
- Platform
- win32 10.0.26200 · x64
- CPU
- 13th Gen Intel(R) Core(TM) i9-13980HX
- Cores
- 32
- Runner
- local
- Generated
- 2026-09-24T18:09:18.839Z
Reproduce
From the Hyperspawn suite monorepo (Coach operator agent, nightly via GitHub Actions):
pnpm --filter @hs/engine engine coach pnpm --filter @hs/engine engine coach --only unitree_g1,booster_t1 --seconds 10
Model files are cached in engine/.cache/coach/models (about 700 MB). Output: this page's JSON plus a dated report under engine/reports/coach.
Next boards
Stand-hold is the floor. Policy evals (velocity tracking error and falls under the same pinned MJCF) rank on the leaderboard once policies arrive with an eval we can rerun; the in-tab walk tests are on each policy page.