Four fixed clips from the actual four-video training manifest. No model predictions and no generated imagery.
Eight sampled frames at 2 FPS; 1008×1008 RGB and 72×72 FP32 camera-Z depth.
Same dataloader, seed and fixed windows as the overfit run. Photometric augmentation and depth dropout are disabled, as in that run.
3D cuboids are smoothed pseudo labels; they are not independent sensor ground truth.
Depth is the model input cache, not full-resolution source depth. Gray/black borders show missing/padded depth.
Video
Objects
Frames with depth
Valid patch fraction
Median depth / m
sav_008220
2
8/8
58.3%
7.872968673706055
sav_006689
3
8/8
58.3%
0.64783775806427
sav_010240
2
8/8
58.3%
0.31044623255729675
sav_002191
1
8/8
58.3%
5.058534622192383
Frame-by-frame inspector
Scrub frames; hover on depth to inspect exact values and the matching RGB patch.
The meter color scale is fixed across frames. Changing the display maximum changes visualization only.
Loading exact dataloader tensors…
Source RGB: before training resize/padding; displayed smallerActual training RGB + manual visible 2D targets (green)Actual training RGB + projected 3D pseudo labels (yellow)Exact 72×72 input depth, nearest-neighbor display; hover for metersTraining RGB + depth; hover depth to mark the same patchPatch valid-pixel coverage: black 0, white 1; hover for exact fraction
Hover over depth or coverage to inspect the exact FP32 values.
Transformed intrinsics, geometry and sample metadata
Video previews · source timing
Each movie uses the sampled source timestamps; eight frames span about 3.5 seconds.
Top row: source RGB, training 2D targets, training 3D targets. Bottom row: depth, RGB/depth overlay, valid-pixel coverage.
Track IDs and segment boundaries are respected; camera-to-world rotations remove estimated camera motion from this comparison.
Raw: original world orientation. Coder: existing per-frame W≤L and yaw folding. Equivalent: minimum local-Y quarter-turn angle, accepting only representations whose dimensions match the previous frame within 10% on every axis.
The equivalent curve is diagnostic, not a relabeling. True motion, near-square ambiguity and pose-estimation errors remain possible.
No representation-jump flags were found in these 35 eligible short-clip transitions using coder angle >60° and equivalent angle <15°. This does not certify the full dataset.
Current labels and training behavior are unchanged. Per-frame canonicalization alone does not establish temporal continuity.
sav_008220 · inspect orientation changes together with size changes; large angles alone are not a bug verdict.sav_006689 · inspect orientation changes together with size changes; large angles alone are not a bug verdict.sav_010240 · inspect orientation changes together with size changes; large angles alone are not a bug verdict.sav_002191 · inspect orientation changes together with size changes; large angles alone are not a bug verdict.