All 9 display videos from this experiment are included below. The latest puppy alignment is first; earlier results are preserved for comparison. Click a video to play and loop it. MP4 files are streamed separately rather than embedded as large GIFs.
Aligned puppy · Person · Multi-asset street · Single-car baseline · Original puppy
Latest version. The independent orthographic display camera was replaced with an RGB-fitted perspective camera. Body placement and heading, head articulation, and the ball position were refitted. Background tracks account for approximate handheld camera motion.
Remaining errors: body proportions, legs, tail and surface appearance are still approximate. Mean body-landmark fitting residual is 14.5 px on a 432 px-wide image, using manual keyframes and interpolated targets. This is a fitting residual, not independent accuracy validation. Camera calibration and metric scale are not recovered ground truth.
Human articulation uses MediaPipe Pose Heavy on RGB: 68 of 72 frames were detected, with four missing frames interpolated. A manually chosen image-space identity prior reduces swaps. Relative depth, standing height and ground contact are assumptions. The mannequin-like geometry and motion are not a photoreal reconstruction.
Only the original black sedan retains RGB-feature-derived motion under a straight-lane prior. The other five vehicle paths and added prop placements are authored synthetic scene completion, not recovered trajectories from the real video. The asset set contains 15 types; boxes cover 24 instances, not every background mesh. Visual appearance remains stylized.
Motion is fitted from manually selected RGB landmarks and optical flow, constrained to a straight lane. A 4.7 m car length supplies an assumed scale. The camera fit is weak and the background and vehicle geometry are simplified. This establishes scene/trajectory export, not accurate monocular metric reconstruction.
Dog translation follows four manually chosen RGB centers after dark-silhouette tracking was contaminated by the baseboard. Limb cycles, head motion and tail motion were authored approximations. Pink-color tracking supplies ball observations; the yellow toy is an authored static prop.
Experiments were generated on 2026-09-10 and published here on 2026-09-17. Inputs were short RGB clips only. No supplied SA-V masks, VGGT depth, SAM3D meshes, FoundationPose outputs or WildTrack3D 3D trajectories were consumed. This is an iterative agent-coded experiment with auxiliary computer vision and manual image-space priors, not an independently audited model-capability test.
Real reference footage: SA-V / Meta, Segment Anything 2, licensed under CC BY 4.0. Reference clips: sav_039074 (car, 0–3 s), sav_041572 (person, 1.5–4.5 s), and sav_025521 (puppy, 0–3 s). Excerpts were resized, annotated and composited alongside synthetic renders; no endorsement is implied.
All nine MP4 files were fully decoded and checked for 72 frames before publication. Media manifest with checksums. No full source videos, scene/model bundles, credentials or internal workspace paths are published in this update.