← Back to project

Text-to-3D Coverage · Gate-1 Pilot — T2I→3D synthesis works

Z-Image vs Qwen-Image → TRELLIS.2 image→3D on 10 coverage-gap concepts · 2026-07-19
text-to-3dsynthetic-datacoveragepilot

1 · What this pilot answers

Verdict: Gate-1 PASSES. Synthetic→3D shows no meaningful domain gap vs real renders; most thin topology reconstructs (gyroid holes, knots, ribs, bond-sticks). Two findings redirect the plan: (1) the bottleneck for exotic concepts is T2I semantics — mandates a post-generation VLM QC gate; (2) one true failure mode: sparse wireframe struts enclosing empty volume get filled solid.

2 · Where the data is missing (measured, not guessed)

Most of the vocabulary has zero or thin coverage — the breadth tier (~283k gen) fills SOC-GC gaps at floor-10; the depth tier (~15k) raises the 48 exotic concepts to ~500 each.
Per-concept counts for the 48 exotic archetypes (log scale). Red = zero, orange = <50. These are exactly the text-to-3D failure modes from the feasibility study.

3 · Gate-1 rows — T2I vs image→3D, concept by concept

Each row: [Z-Image | Z→3D | Qwen | Q→3D | real render | real→3D]. Captions carry the verdicts.

① MOST IMPORTANT ROW — the bottleneck is T2I semantics, not reconstruction. Z-Image collapses 'mobius' to a wavy sheet (no loop, no twist); Qwen gets the half-twist ribbon but not full closure; image→3D faithfully reconstructs whatever it is given. Bonus finding: our library's top-scored 'mobius' (col 5) is actually a garbage text-block asset — the 12 existing ones are low quality.
Zero-coverage concept. Z-Image draws a plain bottle (semantic miss); Qwen draws the neck looping back through the wall — much closer. Both reconstruct cleanly to 3D.
Zero-coverage + hardest topology. Both T2I get porous structures and image→3D PRESERVES the through-holes — the strongest positive signal for the thin tier.
T2I semantics drift (Z-Image: two linked rings; Qwen: celtic-knot cross) but every variant reconstructs with topology intact. real→3D (col 6) is a near-perfect trefoil.
Z-Image degrades to a plain torus; Qwen draws a woven wreath — and image→3D reconstructs even the woven strands. All pass.
THE ONE SYSTEMATIC FAILURE — sparse thin struts enclosing empty volume get reconstructed as a solid/glassy cube (interior filled). Note the dense real building-frame (col 5-6) survives; the failure mode is specifically sparse-strut + large-cavity. Avoid or reformulate this pattern in generation.
Zero-coverage compact control. Both T2I are excellent, both reconstructions keep the cortical folds (see normal maps). Compact tier fully de-risked.
All six columns good, including real→3D. Compact anatomy is safe to generate at scale.
Ball-and-stick with thin bond cylinders reconstructs correctly — thin struts per se are fine (contrast with wireframe row: the issue is enclosed cavities, not thinness).
Full skeleton incl. thin ribs reconstructs from both T2I models and from the real render. Pass.

4 · Decisions locked by this pilot

Coverage: CLIP-matched 1.43M captions vs 37,485 concepts. Pilot: 60 T2I images (2 models × 10 concepts × 3 seeds) + 10 real controls through TRELLIS.2-4B image→3D. Full plan: BLIP3o/docs/TEXT_TO_3D_COVERAGE_PLAN.md