← Back to project

WildTrack3D · Synthetic scene video gallery

Cars, people and a puppy · latest alignment plus preserved baselines · published 2026-09-17
9 videosRGB-guidedNo physicsApproximate geometryNot photoreal

Scope and reading guide

Trajectory export works; photorealism does not. These are iterative, RGB-guided procedural scene experiments, not WildTrack3D model predictions, real-video 3D ground truth, or a verified one-shot model benchmark. No physics simulation is used.

All 9 display videos from this experiment are included below. The latest puppy alignment is first; earlier results are preserved for comparison. Click a video to play and loop it. MP4 files are streamed separately rather than embedded as large GIFs.

Aligned puppy · Person · Multi-asset street · Single-car baseline · Original puppy

1 · Puppy — updated perspective alignment

Latest version. The independent orthographic display camera was replaced with an RGB-fitted perspective camera. Body placement and heading, head articulation, and the ball position were refitted. Background tracks account for approximate handheld camera motion.

Left: real SA-V reference. Right: updated synthetic perspective fit, without boxes so remaining visual differences are visible. Download MP4 · 0.41 MB · 3 s / 72 frames
Show fitted 3D boxes
Boxes are recomputed bounds of the synthetic meshes, not ground-truth boxes in the real clip. Download MP4 · 0.52 MB · 3 s / 72 frames
Show before / after comparison
Left: real reference. Middle: previous independent novel view. Right: updated perspective fit. Download MP4 · 0.52 MB · 3 s / 72 frames

Remaining errors: body proportions, legs, tail and surface appearance are still approximate. Mean body-landmark fitting residual is 14.5 px on a 432 px-wide image, using manual keyframes and interpolated targets. This is a fitting residual, not independent accuracy validation. Camera calibration and metric scale are not recovered ground truth.

2 · Person and renovation props

Reference worker vs a generic procedural human, six paint buckets and an office chair. The generated camera is a separate novel view, not camera-aligned reconstruction. Download MP4 · 0.56 MB · 3 s / 72 frames
Show generated RGB without boxes
Synthetic RGB only. Eight tracked scene entities; other people in the source are omitted. Download MP4 · 0.08 MB · 3 s / 72 frames

Human articulation uses MediaPipe Pose Heavy on RGB: 68 of 72 frames were detected, with four missing frames interpolated. A manually chosen image-space identity prior reduces swaps. Relative depth, standing height and ground contact are assumptions. The mannequin-like geometry and motion are not a photoreal reconstruction.

3 · Multi-asset street

Real reference, synthetic scene, generated box tracks and an elevated overview. Six moving vehicles plus eighteen catalogued static prop instances. Download MP4 · 2.04 MB · 3 s / 72 frames

Only the original black sedan retains RGB-feature-derived motion under a straight-lane prior. The other five vehicle paths and added prop placements are authored synthetic scene completion, not recovered trajectories from the real video. The asset set contains 15 types; boxes cover 24 instances, not every background mesh. Visual appearance remains stylized.

4 · Single-car baseline

The generated box shown on the real video is a projection of the fitted synthetic car; it is not SA-V ground truth or a WildTrack3D prediction. Download MP4 · 2.25 MB · 3 s / 72 frames

Motion is fitted from manually selected RGB landmarks and optical flow, constrained to a straight lane. A 4.7 m car length supplies an assumed scale. The camera fit is weak and the background and vehicle geometry are simplified. This establishes scene/trajectory export, not accurate monocular metric reconstruction.

5 · Puppy — preserved original, unaligned view

Earlier result. Right-hand view uses an independently authored orthographic camera, so its viewpoint does not match the reference. Superseded for viewpoint comparison by Section 1. Download MP4 · 0.54 MB · 3 s / 72 frames
Show original generated RGB without boxes
Preserved original synthetic scene before perspective alignment. Download MP4 · 0.09 MB · 3 s / 72 frames

Dog translation follows four manually chosen RGB centers after dark-silhouette tracking was contaminated by the baseboard. Limb cycles, head motion and tail motion were authored approximations. Pink-color tracking supplies ball observations; the yellow toy is an authored static prop.

Provenance, limitations and source credit

Experiments were generated on 2026-09-10 and published here on 2026-09-17. Inputs were short RGB clips only. No supplied SA-V masks, VGGT depth, SAM3D meshes, FoundationPose outputs or WildTrack3D 3D trajectories were consumed. This is an iterative agent-coded experiment with auxiliary computer vision and manual image-space priors, not an independently audited model-capability test.

Real reference footage: SA-V / Meta, Segment Anything 2, licensed under CC BY 4.0. Reference clips: sav_039074 (car, 0–3 s), sav_041572 (person, 1.5–4.5 s), and sav_025521 (puppy, 0–3 s). Excerpts were resized, annotated and composited alongside synthetic renders; no endorsement is implied.

All nine MP4 files were fully decoded and checked for 72 frames before publication. Media manifest with checksums. No full source videos, scene/model bundles, credentials or internal workspace paths are published in this update.

Research visualization only. Accurate bounds of a synthetic mesh do not imply accurate 3D labels for the real scene.