Text-to-3D Coverage · Gate-1 Pilot — T2I→3D synthesis works
Z-Image vs Qwen-Image → TRELLIS.2 image→3D on 10 coverage-gap concepts · 2026-07-19
text-to-3dsynthetic-datacoveragepilot
1 · What this pilot answers
Question: before spending the 500k synthetic budget, does the T2I→image→3D chain actually work — especially for thin/topological concepts (mobius, gyroid, knots) that text-to-3D currently fails on?
Setup: 10 coverage-gap concepts × 3 seeds × 2 T2I models (Z-Image-Turbo 9-step, Qwen-Image 40-step), each image through vanilla TRELLIS.2-4B image→3D; plus a real library render of the same concept as control.
Columns: [Z-Image | its 3D | Qwen | its 3D | real render | its 3D]. Each 3D cell shows shaded + normals + occupancy strips.
Verdict: Gate-1 PASSES. Synthetic→3D shows no meaningful domain gap vs real renders; most thin topology reconstructs (gyroid holes, knots, ribs, bond-sticks). Two findings redirect the plan: (1) the bottleneck for exotic concepts is T2I semantics — mandates a post-generation VLM QC gate; (2) one true failure mode: sparse wireframe struts enclosing empty volume get filled solid.
2 · Where the data is missing (measured, not guessed)
All 1,433,660 asset captions matched (CLIP text-embedding, not keyword grep) to a 37,485-concept taxonomy = SOC-GC 36,281 (weikaih's CVPR'26 vocabulary) + LVIS 1,156 + 48 hand-authored exotic concepts.
Pool is a game/film library: top-5 categories = 70.8%. 23k of 36k SOC-GC concepts have zero assets.
Exotic tier: 4 concepts at literal zero (klein bottle, gyroid, human brain, generative-art), mobius=12, helicoid=6, neuron=6.
Most of the vocabulary has zero or thin coverage — the breadth tier (~283k gen) fills SOC-GC gaps at floor-10; the depth tier (~15k) raises the 48 exotic concepts to ~500 each.Per-concept counts for the 48 exotic archetypes (log scale). Red = zero, orange = <50. These are exactly the text-to-3D failure modes from the feasibility study.
3 · Gate-1 rows — T2I vs image→3D, concept by concept
Each row: [Z-Image | Z→3D | Qwen | Q→3D | real render | real→3D]. Captions carry the verdicts.
① MOST IMPORTANT ROW — the bottleneck is T2I semantics, not reconstruction. Z-Image collapses 'mobius' to a wavy sheet (no loop, no twist); Qwen gets the half-twist ribbon but not full closure; image→3D faithfully reconstructs whatever it is given. Bonus finding: our library's top-scored 'mobius' (col 5) is actually a garbage text-block asset — the 12 existing ones are low quality.Zero-coverage concept. Z-Image draws a plain bottle (semantic miss); Qwen draws the neck looping back through the wall — much closer. Both reconstruct cleanly to 3D.Zero-coverage + hardest topology. Both T2I get porous structures and image→3D PRESERVES the through-holes — the strongest positive signal for the thin tier.T2I semantics drift (Z-Image: two linked rings; Qwen: celtic-knot cross) but every variant reconstructs with topology intact. real→3D (col 6) is a near-perfect trefoil.Z-Image degrades to a plain torus; Qwen draws a woven wreath — and image→3D reconstructs even the woven strands. All pass.THE ONE SYSTEMATIC FAILURE — sparse thin struts enclosing empty volume get reconstructed as a solid/glassy cube (interior filled). Note the dense real building-frame (col 5-6) survives; the failure mode is specifically sparse-strut + large-cavity. Avoid or reformulate this pattern in generation.Zero-coverage compact control. Both T2I are excellent, both reconstructions keep the cortical folds (see normal maps). Compact tier fully de-risked.All six columns good, including real→3D. Compact anatomy is safe to generate at scale.Ball-and-stick with thin bond cylinders reconstructs correctly — thin struts per se are fine (contrast with wireframe row: the issue is enclosed cavities, not thinness).Full skeleton incl. thin ribs reconstructs from both T2I models and from the real render. Pass.
Prompt style: pure white background, single object, centered (switched from light-grey per weikaih).
New mandatory stage — VLM QC gate: generated image must pass a concept-adherence check (Qwen3.6 judge infra) before entering image→3D. Without it, mislabeled 'wavy sheets called mobius' would poison the training signal.
Wireframe-cavity class: excluded or reformulated (dense-strut variants) — the one verified representation limit.
Thin tier (5,802 gen budget) unlocked — gyroid/knot/mobius class reconstructs; combined with the earlier shape-VAE roundtrip (topology survives our own encode→decode), both layers of the 'representation ceiling' worry are cleared.
Coverage: CLIP-matched 1.43M captions vs 37,485 concepts. Pilot: 60 T2I images (2 models × 10 concepts × 3 seeds) + 10 real controls through TRELLIS.2-4B image→3D. Full plan: BLIP3o/docs/TEXT_TO_3D_COVERAGE_PLAN.md