← Back to project

Interactive 3D — one model, three conditionings

Drag to rotate · scroll to zoom · studio lighting, PBR materials. Every column below comes from the same S3 tri-modal checkpoint family at 35k cumulative; the only thing that changes between I1, IM and T is what the model is conditioned on. Each runs the full cascade with zero GT structure — conditioning → SS flow (64³ occupancy) → shape flow on the generated coords → shape-conditioned texture flow on the generated shape — at an identical seed, so differences are conditioning and not sampling noise. All 14 assets are held out and were never trained on.

I1 · mean SS-IoU@64
0.408
2074 cond tokens · weight 0.5 · 0.363 clean
IM · mean SS-IoU@64
0.371
1897 cond tokens · weight 0.3 · 0.323 clean
T · mean SS-IoU@64
0.202
57 cond tokens · weight 0.2 · 0.141 clean
held-out assets
14
never trained · identical seed 0 · 1 memorised
What changed. This page used to compare GT / TRELLIS.2 / qwen-only / qwen+dino — the axis from when the open question was whether VLM hidden states could condition a 3D flow at all. That question is settled, and the checkpoints that answered it no longer reflect the architecture. The comparison now runs along the axis that is actually open: one model, three conditionings, and how far the cheaper ones (4 views, text alone) fall behind a single clean image.

GT vs I1 vs IM vs T — full cascade

Read the meshes, not the IoU. The badge is the SS stage scored against ground-truth occupancy at 64³, and it is genuinely useful for ranking the three conditionings against each other — but it is a voxel-overlap score, so it collapses on thin structures even when the object is right. The orbital satellite below scores 0.05 and is unmistakably the same satellite; a compact object that is slightly the wrong size scores worse than a blob of the right bulk. Use it to compare columns within a row, not to judge whether a mesh is good. The text column has the strongest claim on this caveat — text→3D is a one-to-many task, so even a perfect answer lands somewhere in a distribution of valid shapes rather than on this one (measured separately in Text-to-3D feasibility: CLIP 0.239 against a GT-render ceiling of 0.248, at a geometric IoU of 0.167).

c0912d54TexVerse

input

conditioning views
IM views [8, 10, 13, 15]
I1 / camera = v005

GT decoded latents

I1 · single image IoU 0.658

IM · 4 views IoU 0.572

T · text only IoU 0.462

T caption A sleek silver luxury sedan featuring a low-slung aerodynamic silhouette, a large front grille, dark tinted windows, and multi-spoke alloy wheels, presented in a clean studio environment.
6fea7992ObjaverseXL_sketchfab also in stages ↓

input

conditioning views
IM views [5, 8, 10, 15]
I1 / camera = v006

GT decoded latents

I1 · single image IoU 0.994

IM · 4 views IoU 0.974

T · text only IoU 0.271

T caption A sleek white sports coupe featuring a low-slung aerodynamic silhouette, a sloping roofline, and a rear engine cover with horizontal vents, resting on a dark rectangular platform.
8d045207SketchfabV1 also in stages ↓

input

conditioning views
IM views [2, 4, 6, 11]
I1 / camera = v009

GT decoded latents

I1 · single image IoU 0.054

IM · 4 views IoU 0.052

T · text only IoU 0.044

T caption A futuristic assault rifle rendered in a uniform matte grey finish, featuring a sleek, angular silhouette with a prominent top-mounted optical sight, a distinct pistol grip, and a long barrel assembly.
0658bbd9SketchfabV1

input

conditioning views
IM views [6, 7, 10, 11]
I1 / camera = v007

GT decoded latents

I1 · single image IoU 0.592

IM · 4 views IoU 0.569

T · text only IoU 0.192

T caption A sleek, white Formula 1 race car with a low-slung, aerodynamic silhouette, featuring a pointed nose cone, exposed wheels, a central cockpit, and a prominent rear wing assembly.
e12ebb10SketchfabV1 also in stages ↓

input

conditioning views
IM views [2, 3, 7, 8]
I1 / camera = v011

GT decoded latents

I1 · single image IoU 0.611

IM · 4 views IoU 0.551

T · text only IoU 0.359

T caption A sleek white convertible sports car with a low-slung, aerodynamic silhouette, featuring a long hood, short rear deck, and an open-top design with visible interior seats and roll bars.
11ca067aObjaverseXL_sketchfab

input

conditioning views
IM views [1, 3, 7, 9]
I1 / camera = v008

GT decoded latents

I1 · single image IoU 0.807

IM · 4 views IoU 0.759

T · text only IoU 0.195

T caption A full-body 3D model of a Star Wars stormtrooper standing upright, featuring a matte white segmented armor suit with a distinct helmet, chest plate, and utility belt, rendered in a neutral pose against a black background.
87e0d7e8ObjaverseXL_github

input

conditioning views
IM views [4, 6, 9, 11]
I1 / camera = v010

GT decoded latents

I1 · single image IoU 0.052

IM · 4 views IoU 0.018

T · text only IoU 0.041

T caption A complex orbital satellite featuring a central cylindrical bus with a large circular dish antenna, flanked by two massive rectangular solar panel arrays extending horizontally from long truss booms, rendered in a uniform matte grey finish.
af33d7c7TexVerse also in stages ↓

input

conditioning views
IM views [7, 10, 12, 13]
I1 / camera = v007

GT decoded latents

I1 · single image IoU 0.363

IM · 4 views IoU 0.080

T · text only IoU 0.023

T caption A vintage portable record player with a rectangular boxy chassis and a hinged lid, featuring a light grey body with a dark front control panel, two circular turntable platters, and a row of knobs and switches.
e088d9fdSketchfabV1

input

conditioning views
IM views [5, 9, 12, 15]
I1 / camera = v008

GT decoded latents

I1 · single image IoU 0.268

IM · 4 views IoU 0.172

T · text only IoU 0.100

T caption A stylized bust sculpture of a female character with a highly reflective chrome finish, featuring a sleek ponytail, oversized sunglasses resting on the forehead, and a smooth, metallic silver surface with blue-tinted accents.
b551fa2aSketchfabV1 also in stages ↓

input

conditioning views
IM views [4, 5, 8, 14]
I1 / camera = v007

GT decoded latents

I1 · single image IoU 0.024

IM · 4 views IoU 0.051

T · text only IoU 0.013

T caption A glossy red rounded square icon featuring a prominent white stylized letter P in the center, rendered as a thick 3D badge with smooth curved edges and a clean, modern aesthetic.
edade2fcSketchfabV1 memorised — reproduced at IoU 1.000 from text alone, so this asset is a training duplicate, not a held-out test

input

conditioning views
IM views [5, 8, 12, 13]
I1 / camera = v011

GT decoded latents

I1 · single image IoU 1.000

IM · 4 views IoU 1.000

T · text only IoU 1.000

T caption A realistic figurine of a standing cow featuring a piebald coat with large brown patches on a white background, small horns, and a visible udder, posed on a flat white base.
9ac88809SketchfabV1

input

conditioning views
IM views [5, 7, 8, 13]
I1 / camera = v009

GT decoded latents

I1 · single image IoU 0.081

IM · 4 views IoU 0.230

T · text only IoU 0.050

T caption A tall, slender, faceted crystal shard with a pointed tip and flat base, featuring a smooth, iridescent surface that shimmers with pearlescent rainbow hues of white, blue, and gold against a dark background.
0635533bSketchfabV1

input

conditioning views
IM views [4, 6, 8, 9]
I1 / camera = v007

GT decoded latents

I1 · single image IoU 0.070

IM · 4 views IoU 0.042

T · text only IoU 0.035

T caption A weathered, moss-covered stone block featuring a broken white column stump and a small round fixture, forming an irregular, low-profile ruin fragment with a rough, textured surface.
af84a920SketchfabV1

input

conditioning views
IM views [2, 5, 7, 13]
I1 / camera = v007

GT decoded latents

I1 · single image IoU 0.143

IM · 4 views IoU 0.125

T · text only IoU 0.044

T caption A short, cylindrical tree stump featuring a rough, deeply furrowed bark texture in shades of grey and brown, with a flat, cut top surface and a slightly wider base resting on the ground.

Stage by stage — where the conditionings diverge

The same cascade, opened up: the SS flow's 64³ occupancy shell (blue), the shape flow decoded on those coordinates (grey), and the textured result. Five assets spanning the full SS-IoU range, not the good ones — the progression is only informative if the failures are in it. Watch the blue shell: by the time structure is wrong there, nothing downstream recovers it, which is why the texture stage can look competent on an object of the wrong shape.

6fea7992ObjaverseXL_sketchfab

input

conditioning views

SS occupancy 64³

what the structure flow generates

shape mesh

shape flow on those coords

textured

tex flow on that shape

I1 · single image

IoU 0.994

IM · 4 views

IoU 0.974

T · text only

IoU 0.271
e12ebb10SketchfabV1

input

conditioning views

SS occupancy 64³

what the structure flow generates

shape mesh

shape flow on those coords

textured

tex flow on that shape

I1 · single image

IoU 0.611

IM · 4 views

IoU 0.551

T · text only

IoU 0.359
af33d7c7TexVerse

input

conditioning views

SS occupancy 64³

what the structure flow generates

shape mesh

shape flow on those coords

textured

tex flow on that shape

I1 · single image

IoU 0.363

IM · 4 views

IoU 0.080

T · text only

IoU 0.023
8d045207SketchfabV1

input

conditioning views

SS occupancy 64³

what the structure flow generates

shape mesh

shape flow on those coords

textured

tex flow on that shape

I1 · single image

IoU 0.054

IM · 4 views

IoU 0.052

T · text only

IoU 0.044
b551fa2aSketchfabV1

input

conditioning views

SS occupancy 64³

what the structure flow generates

shape mesh

shape flow on those coords

textured

tex flow on that shape

I1 · single image

IoU 0.024

IM · 4 views

IoU 0.051

T · text only

IoU 0.013

Provenance

stagecheckpointstate
SS (structure)s3_ss_40k/checkpoint-32000 35k cumulative (32k this run)
shape SLATs3_shape_40k/checkpoint-37000 40k cumulative (37k this run)
texture SLATs3_tex_40k/checkpoint-37000 40k cumulative (37k this run)
I1 / T conditioningv22_heldout v2.2 hidden + DINOv3 (I1); qwen-only (T)
IM conditioningv22_heldout_im4r weighted 4-of-16 views — the distribution IM trained on
SS sampler{"steps": 12, "guidance_strength": 7.5, "guidance_rescale": 0.7, "guidance_interval": [0.6, 1.0], "rescale_t": 5.0} same for all three

Inference mirrors training: cond_adapter=mlp, view_embed_mode=sincos at VIEW_EMBED_SCALE=0.2 (verified: the dino_view_embed table baked into all three checkpoints has row L2 4.525), fuse_dino=True, no position stamp. IM applies the view code to both the DINO and the qwen segment, using the cache's fixed 292-token layout (four 64-token image blocks at 10/76/142/208). T never enters the fusion branch: no DINO segment, no view embedding. Meshes are welded, decimated, re-atlased and re-encoded with KHR_mesh_quantization for the web (56 textured meshes 424→22.0 MB, silhouette IoU 0.942–0.999; 30 stage meshes 487→10.2 MB, silhouette IoU 0.919–1.000). Total committed assets 31.6 MB.