Texture-GT-v2 · Protocol Repair and Pilot20 Leaderboard
WilD3DGen image-to-3D evaluation · 2026-08-17
image-to-3DtextureToys4KLumiTexaudit
TL;DR
The repaired 20-case × 3-model matrix passes every protocol gate. S3-T50 leads all eight primary paired and dataset-level base-color metrics.
The old scorer call was upstream-faithful, but its asset contract was not: models used different geometry inputs and the canonical rebake mesh came from S3's decoded representation.
This is a controlled intrinsic/base-color track, not an end-to-end geometry score and not a metallic/roughness or relighting benchmark.
1 · What changed
Contract
Legacy proxy
Texture-GT-v2
Canonical geometry
S3 decoded GT latent
Raw Toys4K released mesh
GT appearance
S3 decoded PBR proxy
Original Toys4K Blender base color
Model input geometry
Mixed by model
One shared GT geometry
Coordinates
S3 adapter missing
Frozen model-level adapters
Distribution metrics
Single pair then mean
Released values + proper dataset-level values
LumiTex scorer code remains unchanged at commit bf297538....
Toys4K source is pinned at commit 6cfc6bd....
No per-case orientation search or metric-aware alignment is allowed.
The repaired path: one input, original GT appearance, identical 12-view cameras, model renders, and a geometry-only silhouette gate before rebaking or scoring.
2 · Hard validation gates
20 unique Toys4K cases × 3 model outputs × 5 released metrics are complete and finite.
All model atlases and GT atlases use the same raw-GT xatlas layout and shared valid mask.
Every Blender-GT view must match the LumiTex raster silhouette at IoU ≥ 0.90.
The lowest observed per-view IoU is 0.919 on the bicycle case; most cases exceed 0.98.
All 20 cases clear the frozen 0.90 gate, ruling out camera, raster-origin, and gross coordinate mismatches before metrics are computed.
3 · Pilot20 results
Model
CLIP-I ↑
LPIPS ↓
PSNR ↑
SSIM ↑
Dataset FID ↓
Dataset CLIP-FID ↓
Dataset CMMD ↓
TRELLIS.2
0.9297
0.2761
13.905
0.5086
166.444
16.536
0.1801
Hunyuan3D-2.1
0.9127
0.3040
12.682
0.4738
180.017
20.078
0.2074
S3-T50
0.9415
0.2037
14.923
0.5526
134.813
13.595
0.1683
S3-T50 ranks first on all eight displayed metrics under the controlled GT-geometry contract.
Dataset FID, CLIP-FID, and CMMD are computed once over each complete 20-image set.
CLIP-I, LPIPS, masked PSNR, and masked SSIM remain paired-image metrics.
S3-T50 leads semantic, perceptual, pixel-level, and distribution-level base-color measurements after the asset contract is made model-neutral.
4 · Visual audit
Gray texels are invalid or unobserved atlas regions and are excluded from masked pixel diagnostics.
Atlas topology is identical across each row, enabling direct texture comparison without geometry confounds.
The grid includes articulated, thin-structure, saturated, and multi-material examples.
Four representative canonical-atlas rows show the original Toys4K target next to all model predictions on exactly the same UV support.
5 · Interpretation boundaries
Promotable claim: controlled base-color texture quality on shared GT geometry for this fixed 20-case pilot.
Not yet promotable: exact LumiTex paper-number reproduction; its paper split contains 133 cases, while this pilot uses 20 deterministic Toys4K categories.
Not measured here: generated geometry, metallic, roughness, normals, lighting response, or full-asset end-to-end quality.
Next scale gate: run the frozen protocol on the full released split without changing adapters, cameras, masks, or aggregation.
Texture-GT-v2 · strict validator passed · 20 cases × 3 models · original Toys4K base color