← Back to project

WilD3DGen v10 — three-tower union DiT, 20k steps

3.9B geo/tex/SS union MMDiT · 32×H200 · 40h · checkpoint-20000 · 2026-08-28
evalattention-probetrellis2-comparison

TL;DR

Texture improved on every single asset across all three tasks — 12/0, p<0.001. tex_std deviation 0.40 → 0.12.
Geometry did not move. geo_mse p = 0.30 / 0.28 / 0.58 across the three tasks. 20,000 steps, 3.9B params, no significant change.
The towers did not learn to ignore each other. They narrowed the cross-modal budget 37–65% while making it 15–30× more precise, forming a spatial chain SS → geo → tex in which each stream reads the co-located token of the one before.
Two official baselines recorded in our eval skill do not reproduce on val200: TRELLIS.2-4B G3 measured 0.231 (recorded 0.091) and G4 1.159 (recorded 1.40). Both were held-out-set specific.

1 · Everything in one view

The towers separated, texture won, geometry stood still.

2 · Do the cross-modal attentions leverage or ignore?

Answer: leverage, selectively. Total mass is the wrong statistic — it fell, while precision rose 15–30×.

3 · The other direction: does slat read SS?

geo reads SS twice as hard as tex does, at every noise level.

4 · Coupled inference: real, learned, and too small to use

Three measurements, one answer: real path, negligible budget.
Every tile is a decoded 64³ occupancy of x̂₀ — the orange strip is the slat noise level SS actually reads at that step.

5 · Three tasks, end to end

v10 @20000.
warm @0, properly matched — its own connectors, not the trained ones.

6 · Two eval-harness bugs found along the way

7 · Open questions

single seed · n=32 occ_iou / n=12 latent / n=16 G3 · val200 + val200_capT · checkpoint-20000