The active system is X2-native at the robot, data, state, action, and rollout layers. The only remaining g1 label is a compatibility key inside old serialized SONIC checkpoints. The current project-local sphere continuation has reached iteration 46.2k on Weka as observed at report build time, but 44k remains the newest sphere checkpoint with completed eval and MuJoCo tests.
| Layer | Authoritative choice | Explicit exclusion |
|---|---|---|
| Retarget | BONES-SMPL -> GMR X2 v4, pelvis/leg scale 0.787, virtual toe target at [0.10, 0, -0.073] m | SOMA and SOMA/GMR grafts are closed; do not revive them |
| Training robot | 29 actuated DOFs, physical head retained, two neck DOFs fixed, X2 wrist_roll endpoints | No toe marker, toe joint, toe mass, or extra policy body slot |
| Corpus | 121,486 qpos clips and matched MotionLib clips at 50 Hz | No v3 blanket FOOT_OFFSET or metadata-only FPS change |
| Training | SONIC PPO/FSQ architecture, X2 asset/gains/order, X2-native data | HoloMotion is optional research, not a replacement justified by the invalid 38% |
| Deployment | Official learned G1 kinematic planner -> online GMR v4 -> X2 SONIC -> MuJoCo | This is not a native learned X2 planner and not a physics-aware planner |
Compared with the previous 0.9 production-like retarget corpus, v4 median horizontal root speed is 0.883x, or 11.7% slower. This is a morphology/calibration change, not an artificial duration change.
| Component | Exact setting |
|---|---|
| Motions | 512 unique frozen v4 motions; 128 environments x four loops; start at frame zero |
| Policy | SONIC im_eval, deterministic act_inference |
| Encoder | Public use_encoder=x2; legacy serialized state-dict key g1 only inside the compatibility resolver |
| Failure terms | root Z error > 0.25 m; ankle/wrist_roll endpoint Z error > 0.25 m; root orientation error > 1 rad |
| Success | Motion reaches timeout without an earlier failure |
| Paper/source correction | foot_pos_xyz=null, removing the release merge's extra 0.2 m full ankle-XYZ term |
| Runtime retained | Seed 0, startup physics DR, observation corruption, and trimesh microterrain; no pushes or reset/motion augmentation |
| Metrics | SMPLSim-equivalent formulas; success-only and all-motion scopes are both reported |
| Checkpoint | Success | Progress | Success MPJPE-L | All MPJPE-L | All MPJPE-G |
|---|---|---|---|---|---|
| Mainline 50k | 480/512 (93.75%) | 95.32% | 44.29 mm | 47.56 mm | 249.64 mm |
| Mainline 26k | 477/512 (93.16%) | 94.54% | 41.90 mm | 45.79 mm | 426.18 mm |
| Mainline 34k | 477/512 (93.16%) | 94.96% | 43.59 mm | 46.89 mm | 372.22 mm |
| Mainline 42k | 476/512 (92.97%) | 94.83% | 42.93 mm | 46.92 mm | 286.69 mm |
| No-rescale 15.9k | 465/512 (90.82%) | 92.49% | 40.70 mm | 45.38 mm | 429.05 mm |
| Anchor-v2 8k | 465/512 (90.82%) | 93.15% | 46.31 mm | 51.08 mm | 299.83 mm |
| G1 warm-start 12.3k | 464/512 (90.62%) | 93.60% | 43.45 mm | 48.04 mm | 501.43 mm |
| v4 scratch 2.2k | 412/512 (80.47%) | 87.13% | 57.39 mm | 63.67 mm | 419.75 mm |
Mainline 50k is the winner by success, average completion progress, and global tracking. The no-rescale checkpoint has lower local MPJPE-L but worse success and much worse global error, so MPJPE-L alone is not a checkpoint selector.
For paper context only, SONIC reports 23.2 mm MPJPE-L in its main result, 27.5/26.3 mm for its 32-GPU FSQ-32-32 ablation, and 23.8/22.5 mm for its 128-GPU robot encoder. X2's 44.29/47.56 mm successful/all-motion values are higher (worse), and the embodiment, data, scale, and test set differ.
| Interval | Reward profile | Reason |
|---|---|---|
| 0 to 10k | X2 reform: anchor weight 2.0; body linear-velocity std 0.3 | Stronger global anchoring and tighter speed tracking for early X2 learning |
| 10k onward | Official: anchor weight 0.5; body linear-velocity std 1.0 | Return to upstream reward weighting while preserving the corrected X2 embodiment and data |
| Sphere 42k onward | Official profile | Contact adaptation only; no lateral reward merged |
The first official-reward continuation eval was 459/512 at 11.1k; 20k reached 467/512 (91.21%), 93.90% progress, 45.52 mm successful MPJPE-L, 49.19 mm all-motion MPJPE-L, and 360.64 mm all-motion MPJPE-G. This newer from-scratch mesh continuation reached 50k on disk; its final primary-protocol evaluation is not asserted here.

This ledger consolidates transcript decisions, project memory, source inspection, and generated artifacts. Superseded files/results are retained only when they are useful forensic evidence.
| Issue | Failure mode | Status | Resolution / boundary |
|---|---|---|---|
| X2 wrist torque transcription | 2.2 instead of 4.8 Nm; wrist control was artificially weak. | Fixed | X2 actuator constants now match URDF/MJCF. |
| G1/X2 joint ordering | Six warm-start joints were not semantically aligned: waist pitch/roll and both wrist yaw/roll pairs. | Fixed | Deterministic permutation at the checkpoint boundary; X2 remains 29 DOF. |
| G1 hand endpoint hardcode | G1 uses wrist_yaw as its terminal link; X2's true endpoint is wrist_roll. | Fixed | Reward, termination, contact subsets, eval metrics, and mass randomization use wrist_roll. |
| Empty reward body selection | G1 wrist patterns selected zero X2 bodies and produced NaN reward reductions on every rank. | Fixed | Configured X2 names plus fail-fast validation for empty subsets. |
| Physically headless X2 assets | An old model removed the head mass while the intended embodiment only fixes neck DOFs. | Deleted | Physical head retained; neck yaw/pitch fixed. Old headless files are not valid inputs. |
| Foot endpoint/ankle offset | v3 treated the ankle as the foot and used a blanket FOOT_OFFSET/world-foot patch. | Replaced | v4 uses calibrated X2 scale and massless, jointless, geometry-free virtual toe targets during GMR only. |
| Stale GMR production config | A copied patch could silently restore the pre-wrist-fix mapping. | Fixed | Canonical v4 config is the source of truth; stale alternatives are excluded. |
| Frame-rate relabeling | Smoothing could lower effective bandwidth without actually resampling, while metadata was relabeled. | Fixed | Frame rate changes require real resampling; v4 GMR/MotionLib output is 50 Hz. |
| Custom 38% evaluator | It checked root XY and full 3D endpoints, then was mislabeled as SONIC success. | Deleted | Official im_eval state machine is used; 38% is invalid as a policy score or framework ceiling. |
| Release eval config merge | OmegaConf retained training foot_pos_xyz, adding an unpublished 0.2 m full ankle-XYZ termination. | Controlled | Primary protocol explicitly sets foot_pos_xyz=null; released-code results remain separately labeled. |
| Mixed encoder eval | Unset use_encoder sampled robot, teleop, and SMPL encoders. | Fixed | Primary eval requests use_encoder=x2 and maps only the serialized key to legacy g1. |
| SMPLSim metric dependency | The external package was unavailable in the eval image. | Fixed | Vendored formula matches SMPLSim b5c0872 exactly in deterministic parity tests. |
| Planner stitching freeze | Clipped old indices were reused, freezing 130 of 299 frames. | Fixed | Absolute timeline stitching produced 0 frozen frames in the 1,099-frame regression. |
| MuJoCo action clipping | Deployment fed unclipped policy actions although Isaac clips them to +/-20 before env.step. | Fixed | The policy output and next last_action observation are clipped identically; false 42k/42.4k fall conclusions were withdrawn. |
| Planner token profile | The runner defaulted to local 16-token long horizon; 0.1 s RUN replans collapsed/skated the G1 reference. | Fixed | Default is official 9/10/11-token behavior. Long is diagnostic only; G1 and X2 references are now recorded together. |
| Data conversion scheduling | A grouped 128-task conversion stopped when one interruptible replica was preempted. | Fixed | Conversion uses independent interruptible jobs; grouped execution remains appropriate for synchronous training. |
A short 10k to 10.2k A/B used 12 original/mirrored forward walks. Control kept the 10k reward; treatment added only lateral-velocity tracking. The hard term (weight 1.0, std 0.1) reduced median endpoint |Y| by about 26% and modestly improved cross-track and lateral-velocity RMS. The soft term (weight 0.5, std 0.2) worsened endpoint and cross-track error.


The original X2 training foot uses one triangle-mesh collision per foot. The project-local treatment uses 12 symmetric spheres per foot: ten sole samples plus two toe-edge samples. Official SONIC G1 training uses seven cylinders per foot; its MotionLib reference asset uses four spheres. Therefore, 12 spheres is a standard primitive-contact strategy and symmetric by construction, but it is not an upstream SONIC default.
| Sphere 44k result | Value | Interpretation |
|---|---|---|
| Released-code eval success | 450/512 (87.89%) | Matching sphere physics, but includes release termination behavior; not directly comparable to the 93.75% primary table |
| Progress | 91.03% | Same released-code sphere run |
| Success / all MPJPE-L | 39.52 / 43.25 mm | Lower local pose error does not establish higher success under another protocol |
| 20 s slow walk | 0 falls; 3.65 m forward | Stable after action-clip fix |
| Direction residual | -0.474 m lateral; -1.30 deg yaw vs reference | Sphere contacts improve some yaw probes but do not eliminate lateral tracking error |

Earlier 42.3k/42.4k sphere runs reported 12 falls and suggested later training had regressed. The MuJoCo runner was missing Isaac's manager-environment action clip. After clipping the policy output to +/-20 before both control and the next last_action observation, the 44k sphere policy completed the matched slow-walk probe with zero falls.
The WASD stack uses SONIC's learned G1 kinematic planner, then online GMR to produce X2 references. A local long profile forced 16 output tokens (about 106 frames at 50 Hz). In RUN mode the planner replans every 0.1 s, so repeatedly feeding back this long horizon collapsed the G1 root height and created skating/floating references. The old run claimed 23.13 m of reference travel in 14 s while the robot barely moved; that result is invalid.
The runner now defaults to the official 9/10/11-token behavior, records G1 and X2 references side by side, and retains long only as a named diagnostic. With the corrected profile, G1 reference root Z stays 0.694-0.789 m, X2 reference root Z stays 0.593-0.676 m, and the 44k sphere policy travels about 6.18 m in the six-second RUN probe with zero falls.
A geometry-controlled high-jump probe measured true collision-surface clearance. All three X2 variants produced zero frames with both feet more than 1 mm above ground: 42k mesh, 42k sphere out-of-distribution, and 10.2k sphere-trained. Peak simultaneous clearance was approximately -0.08 mm, +0.014 mm, and +0.001 mm respectively.
These clips visualize the v4 reference layer, not policy rollouts. They demonstrate the breadth passed into MotionLib and expose residual kinematic artifacts before reinforcement learning.
This side-by-side keeps the GMR walking reference and IsaacLab environment aligned while comparing the 10k X2-anchor boundary against the 20k official-reward continuation. The official continuation improves local tracking over time but relaxes direct global velocity pressure; longitudinal speed alone is not the success criterion.
| Open item | What is known | Required gate |
|---|---|---|
| 46.2k sphere continuation | Checkpoint exists on Weka; training uses the official reward profile | Export and run the same primary 512 protocol plus matched MuJoCo direction/jump probes |
| Sphere production choice | 44k is stable in corrected MuJoCo and can improve yaw; primary-protocol ranking is missing | Do not replace mesh baseline until primary success, drift, and jump gates all pass |
| Jumping | No tested X2 variant achieves >1 mm two-foot flight | Audit aerial data/contact labels and controlled reward/actuation ablations |
| Direction bias | Hard lateral reward helps aggregates but has outliers; spheres do not fully remove residual | Implement X2 physics-side symmetry/flip maps, then repeat paired original/mirrored tests |
| Online reference | Official planner profile is stable; GMR remains kinematic with no contact lock | Add contact-aware postprocessing or train a native X2 planner if interactive quality requires it |
| Armature | Uniform 0.03 is still an unverified placeholder | Obtain vendor motor/reflected-inertia ground truth before sim-to-real claims |
| Mass assets | GMR MJCF 43.4749 kg; Isaac URDF 41.9665 kg; deployment MuJoCo 41.9665 kg after pelvis fix | Keep the kinematic GMR asset separate from the dynamics authority |
Mirrored corpus doubling is data symmetry, not simulator symmetry augmentation. HoloMotion remains a bounded future comparison. SOMA is closed. The public report contains no credentials, private tokens, or raw chat excerpts.
| Claim | Repository evidence |
|---|---|
| Primary eval contract and scorecard | docs/x2_sonic_official_eval_20260725.md; logs_rl_beaker/eval_v4/*robot_paper* |
| Upstream/X2 semantic audit | docs/sonic_upstream_audit_20260724.md |
| Retarget and embodiment rules | CLAUDE.md; x2_sonic_training/gear_sonic_x2/PLANNING.md; project memory |
| Sphere 44k eval | logs_rl_beaker/eval_v4/sphere-44000_v4_512_metrics/metrics_eval.json |
| Corrected MuJoCo walk | artifacts/x2_mujoco_regression/sphere_044000_clip20_metrics.json |
| Planner profile diagnostic | artifacts/x2_sphere_44k_run/run_mode3_speed1p8_official_diag.npz; scripts/run_x2_sonic_mujoco.py |
| Jump flight metric | artifacts/x2_jump_checkpoint_ablation/high_jump_flight_metrics.json |
| Lateral reward A/B | artifacts/lateral_ablation_10200/; sonic_x2_lateral_vel*.yaml |
| Foot geometry | gear_sonic_x2/visualize_foot_collisions.py; parsed URDF/MJCF assets |
Frozen eval manifest SHA256: 05c108fdf04bb98b0dd4b1a6f99cc29776437d70c43b7b419769a9829e3d2a11. This page is the current status authority through 2026-07-31; older Hub reports remain historical snapshots and may contain explicitly superseded conclusions.