← Back to project
Interactive 3D — one model, three conditionings
Drag to rotate · scroll to zoom · studio lighting, PBR materials.
Every column below comes from the same S3 tri-modal checkpoint family at
35k cumulative; the only thing that changes between I1, IM and T is what the model is
conditioned on. Each runs the full cascade with zero GT structure — conditioning →
SS flow (64³ occupancy) → shape flow on the generated coords → shape-conditioned
texture flow on the generated shape — at an identical seed, so differences are
conditioning and not sampling noise. All 14 assets are held out and were never
trained on.
I1 · mean SS-IoU@64
0.408
2074 cond tokens · weight 0.5 · 0.363 clean
IM · mean SS-IoU@64
0.371
1897 cond tokens · weight 0.3 · 0.323 clean
T · mean SS-IoU@64
0.202
57 cond tokens · weight 0.2 · 0.141 clean
held-out assets
14
never trained · identical seed 0 · 1 memorised
What changed. This page used to compare GT / TRELLIS.2 /
qwen-only / qwen+dino — the axis from when the open question was whether VLM hidden states
could condition a 3D flow at all. That question is settled, and the checkpoints that
answered it no longer reflect the architecture. The comparison now runs along the axis that
is actually open: one model, three conditionings, and how far the cheaper ones
(4 views, text alone) fall behind a single clean image.
GT vs I1 vs IM vs T — full cascade
Read the meshes, not the IoU. The badge is the SS stage scored against ground-truth
occupancy at 64³, and it is genuinely useful for ranking the three conditionings against each
other — but it is a voxel-overlap score, so it collapses on thin structures even when the
object is right. The orbital satellite below scores 0.05 and is unmistakably the same
satellite; a compact object that is slightly the wrong size scores worse than a blob of the
right bulk. Use it to compare columns within a row, not to judge whether a mesh is good.
The text column has the strongest claim on this caveat — text→3D is a one-to-many task, so
even a perfect answer lands somewhere in a distribution of valid shapes rather than on
this one (measured separately in
Text-to-3D feasibility: CLIP 0.239 against a
GT-render ceiling of 0.248, at a geometric IoU of 0.167).
c0912d54TexVerse
input

IM views [8, 10, 13, 15]
I1 / camera = v005
GT decoded latents
I1 · single image IoU 0.658
IM · 4 views IoU 0.572
T · text only IoU 0.462
T caption A sleek silver luxury sedan featuring a low-slung aerodynamic silhouette, a large front grille, dark tinted windows, and multi-spoke alloy wheels, presented in a clean studio environment.
6fea7992ObjaverseXL_sketchfab also in stages ↓
input

IM views [5, 8, 10, 15]
I1 / camera = v006
GT decoded latents
I1 · single image IoU 0.994
IM · 4 views IoU 0.974
T · text only IoU 0.271
T caption A sleek white sports coupe featuring a low-slung aerodynamic silhouette, a sloping roofline, and a rear engine cover with horizontal vents, resting on a dark rectangular platform.
8d045207SketchfabV1 also in stages ↓
input

IM views [2, 4, 6, 11]
I1 / camera = v009
GT decoded latents
I1 · single image IoU 0.054
IM · 4 views IoU 0.052
T · text only IoU 0.044
T caption A futuristic assault rifle rendered in a uniform matte grey finish, featuring a sleek, angular silhouette with a prominent top-mounted optical sight, a distinct pistol grip, and a long barrel assembly.
0658bbd9SketchfabV1
input

IM views [6, 7, 10, 11]
I1 / camera = v007
GT decoded latents
I1 · single image IoU 0.592
IM · 4 views IoU 0.569
T · text only IoU 0.192
T caption A sleek, white Formula 1 race car with a low-slung, aerodynamic silhouette, featuring a pointed nose cone, exposed wheels, a central cockpit, and a prominent rear wing assembly.
e12ebb10SketchfabV1 also in stages ↓
input

IM views [2, 3, 7, 8]
I1 / camera = v011
GT decoded latents
I1 · single image IoU 0.611
IM · 4 views IoU 0.551
T · text only IoU 0.359
T caption A sleek white convertible sports car with a low-slung, aerodynamic silhouette, featuring a long hood, short rear deck, and an open-top design with visible interior seats and roll bars.
11ca067aObjaverseXL_sketchfab
input

IM views [1, 3, 7, 9]
I1 / camera = v008
GT decoded latents
I1 · single image IoU 0.807
IM · 4 views IoU 0.759
T · text only IoU 0.195
T caption A full-body 3D model of a Star Wars stormtrooper standing upright, featuring a matte white segmented armor suit with a distinct helmet, chest plate, and utility belt, rendered in a neutral pose against a black background.
87e0d7e8ObjaverseXL_github
input

IM views [4, 6, 9, 11]
I1 / camera = v010
GT decoded latents
I1 · single image IoU 0.052
IM · 4 views IoU 0.018
T · text only IoU 0.041
T caption A complex orbital satellite featuring a central cylindrical bus with a large circular dish antenna, flanked by two massive rectangular solar panel arrays extending horizontally from long truss booms, rendered in a uniform matte grey finish.
af33d7c7TexVerse also in stages ↓
input

IM views [7, 10, 12, 13]
I1 / camera = v007
GT decoded latents
I1 · single image IoU 0.363
IM · 4 views IoU 0.080
T · text only IoU 0.023
T caption A vintage portable record player with a rectangular boxy chassis and a hinged lid, featuring a light grey body with a dark front control panel, two circular turntable platters, and a row of knobs and switches.
e088d9fdSketchfabV1
input

IM views [5, 9, 12, 15]
I1 / camera = v008
GT decoded latents
I1 · single image IoU 0.268
IM · 4 views IoU 0.172
T · text only IoU 0.100
T caption A stylized bust sculpture of a female character with a highly reflective chrome finish, featuring a sleek ponytail, oversized sunglasses resting on the forehead, and a smooth, metallic silver surface with blue-tinted accents.
b551fa2aSketchfabV1 also in stages ↓
input

IM views [4, 5, 8, 14]
I1 / camera = v007
GT decoded latents
I1 · single image IoU 0.024
IM · 4 views IoU 0.051
T · text only IoU 0.013
T caption A glossy red rounded square icon featuring a prominent white stylized letter P in the center, rendered as a thick 3D badge with smooth curved edges and a clean, modern aesthetic.
edade2fcSketchfabV1 memorised — reproduced at IoU 1.000 from text alone, so this asset is a training duplicate, not a held-out test
input

IM views [5, 8, 12, 13]
I1 / camera = v011
GT decoded latents
I1 · single image IoU 1.000
IM · 4 views IoU 1.000
T · text only IoU 1.000
T caption A realistic figurine of a standing cow featuring a piebald coat with large brown patches on a white background, small horns, and a visible udder, posed on a flat white base.
9ac88809SketchfabV1
input

IM views [5, 7, 8, 13]
I1 / camera = v009
GT decoded latents
I1 · single image IoU 0.081
IM · 4 views IoU 0.230
T · text only IoU 0.050
T caption A tall, slender, faceted crystal shard with a pointed tip and flat base, featuring a smooth, iridescent surface that shimmers with pearlescent rainbow hues of white, blue, and gold against a dark background.
0635533bSketchfabV1
input

IM views [4, 6, 8, 9]
I1 / camera = v007
GT decoded latents
I1 · single image IoU 0.070
IM · 4 views IoU 0.042
T · text only IoU 0.035
T caption A weathered, moss-covered stone block featuring a broken white column stump and a small round fixture, forming an irregular, low-profile ruin fragment with a rough, textured surface.
af84a920SketchfabV1
input

IM views [2, 5, 7, 13]
I1 / camera = v007
GT decoded latents
I1 · single image IoU 0.143
IM · 4 views IoU 0.125
T · text only IoU 0.044
T caption A short, cylindrical tree stump featuring a rough, deeply furrowed bark texture in shades of grey and brown, with a flat, cut top surface and a slightly wider base resting on the ground.
Stage by stage — where the conditionings diverge
The same cascade, opened up: the SS flow's 64³ occupancy shell (blue), the shape flow
decoded on those coordinates (grey), and the textured result. Five assets spanning the full
SS-IoU range, not the good ones — the progression is only informative if the failures are in
it. Watch the blue shell: by the time structure is wrong there, nothing downstream recovers
it, which is why the texture stage can look competent on an object of the wrong shape.
6fea7992ObjaverseXL_sketchfab
input

SS occupancy 64³
what the structure flow generates
shape mesh
shape flow on those coords
textured
tex flow on that shape
I1 · single image
IoU 0.994
IM · 4 views
IoU 0.974
T · text only
IoU 0.271
e12ebb10SketchfabV1
input

SS occupancy 64³
what the structure flow generates
shape mesh
shape flow on those coords
textured
tex flow on that shape
I1 · single image
IoU 0.611
IM · 4 views
IoU 0.551
T · text only
IoU 0.359
af33d7c7TexVerse
input

SS occupancy 64³
what the structure flow generates
shape mesh
shape flow on those coords
textured
tex flow on that shape
I1 · single image
IoU 0.363
IM · 4 views
IoU 0.080
T · text only
IoU 0.023
8d045207SketchfabV1
input

SS occupancy 64³
what the structure flow generates
shape mesh
shape flow on those coords
textured
tex flow on that shape
I1 · single image
IoU 0.054
IM · 4 views
IoU 0.052
T · text only
IoU 0.044
b551fa2aSketchfabV1
input

SS occupancy 64³
what the structure flow generates
shape mesh
shape flow on those coords
textured
tex flow on that shape
I1 · single image
IoU 0.024
IM · 4 views
IoU 0.051
T · text only
IoU 0.013
Provenance
| stage | checkpoint | state |
| SS (structure) | s3_ss_40k/checkpoint-32000 |
35k cumulative (32k this run) |
| shape SLAT | s3_shape_40k/checkpoint-37000 |
40k cumulative (37k this run) |
| texture SLAT | s3_tex_40k/checkpoint-37000 |
40k cumulative (37k this run) |
| I1 / T conditioning | v22_heldout |
v2.2 hidden + DINOv3 (I1); qwen-only (T) |
| IM conditioning | v22_heldout_im4r |
weighted 4-of-16 views — the distribution IM trained on |
| SS sampler | {"steps": 12, "guidance_strength": 7.5, "guidance_rescale": 0.7, "guidance_interval": [0.6, 1.0], "rescale_t": 5.0} |
same for all three |
Inference mirrors training: cond_adapter=mlp,
view_embed_mode=sincos at VIEW_EMBED_SCALE=0.2 (verified: the
dino_view_embed table baked into all three checkpoints has row L2 4.525),
fuse_dino=True, no position stamp. IM applies the view code to both the DINO and
the qwen segment, using the cache's fixed 292-token layout (four 64-token image blocks at
10/76/142/208). T never enters the fusion branch: no DINO segment, no view embedding.
Meshes are welded, decimated, re-atlased and re-encoded with KHR_mesh_quantization for the web (56 textured meshes 424→22.0 MB, silhouette IoU 0.942–0.999; 30 stage meshes 487→10.2 MB, silhouette IoU 0.919–1.000). Total committed assets 31.6 MB.