official recipe · warm-start basis64_5000 · iter 6,593/20,000 · 2026-07-11
traininglivehumanoid8xGPU
1 · Run at a glance
LIVE — iteration 6,593 / 20,000 (33.0%) · mean reward 2.133 (from 2.079) · joint error 0.208 rad (from 0.227)
field
value
Hardware
8× GPU, single node (AI2 Beaker, isaac-sim 5.1.0 image)
Recipe
Official — zero reward overrides (the validated recipe)
Warm start
basis64_5000.pt (prior champion, cadence 9/14)
Corpus
61,677 unique clips (full BONES-SEED, quality-gated); auto-upgrades to 122,572 with mirrors on any restart
Parallel envs
8 × 4096 = 32,768 (16× the prior champion's batch)
Speed / ETA
~9.1 s/iter → ~35 h remaining to 20k
Checkpoints
every 250 iters → weka logs_rl_beaker/run_b64full_8gpu/; preemption-safe resume
2 · Live training curves
Parsed from the live job log. Iteration counter resumes from the basis64 seed at 5,000 (dashed line); everything to its right is new full-corpus training. Thin line = raw per-iteration, thick = 15-iter moving average.
reward climbing, tracking errors falling — healthy early training on 61.7k diverse clips
3 · What to expect
Early terminations are high (ee_body_pos ~0.5) because the policy is meeting dances, lifts and gestures it has never seen; the adaptive sampler is up-weighting hard clips (effective bins 361 → 1,592).
Mid-run checkpoint will be probed with the core-16 cadence gate — the same ruler that scored basis64_5000 at 9/14 — to measure what full-corpus diversity buys.
Data pipeline is closed: 131k BONES-SEED → GMR retarget (validated to 1.28° vs production v3) → physics QA (6.3% flips removed) → 122,572-clip balanced corpus incl. mirrors.
auto-generated from the live Beaker job log · updates on re-publish