← Back to project

Texture-GT-v2 · Protocol Repair and Pilot20 Leaderboard

WilD3DGen image-to-3D evaluation · 2026-08-17
image-to-3DtextureToys4KLumiTexaudit

TL;DR

The repaired 20-case × 3-model matrix passes every protocol gate. S3-T50 leads all eight primary paired and dataset-level base-color metrics.
The old scorer call was upstream-faithful, but its asset contract was not: models used different geometry inputs and the canonical rebake mesh came from S3's decoded representation.
This is a controlled intrinsic/base-color track, not an end-to-end geometry score and not a metallic/roughness or relighting benchmark.

1 · What changed

ContractLegacy proxyTexture-GT-v2
Canonical geometryS3 decoded GT latentRaw Toys4K released mesh
GT appearanceS3 decoded PBR proxyOriginal Toys4K Blender base color
Model input geometryMixed by modelOne shared GT geometry
CoordinatesS3 adapter missingFrozen model-level adapters
Distribution metricsSingle pair then meanReleased values + proper dataset-level values
The repaired path: one input, original GT appearance, identical 12-view cameras, model renders, and a geometry-only silhouette gate before rebaking or scoring.

2 · Hard validation gates

All 20 cases clear the frozen 0.90 gate, ruling out camera, raster-origin, and gross coordinate mismatches before metrics are computed.

3 · Pilot20 results

ModelCLIP-I ↑LPIPS ↓PSNR ↑SSIM ↑Dataset FID ↓Dataset CLIP-FID ↓Dataset CMMD ↓
TRELLIS.20.92970.276113.9050.5086166.44416.5360.1801
Hunyuan3D-2.10.91270.304012.6820.4738180.01720.0780.2074
S3-T500.94150.203714.9230.5526134.81313.5950.1683
S3-T50 leads semantic, perceptual, pixel-level, and distribution-level base-color measurements after the asset contract is made model-neutral.

4 · Visual audit

Four representative canonical-atlas rows show the original Toys4K target next to all model predictions on exactly the same UV support.

5 · Interpretation boundaries

Texture-GT-v2 · strict validator passed · 20 cases × 3 models · original Toys4K base color