Merged suite specification, metric provenance, repository integration, and implementation gates · updated 2026-08-11
| Track | Full split | Required output | Primary question |
| Paired-GT final image-to-3D | Fixed Toys4K pool: 3,229 public assets; deterministic source views | Final generated mesh/GLB | How accurately does the complete model recover known 3D geometry from one image? |
| Open-image final generation | 100 fixed public/licensed AI images in the TRELLIS.2 style | Final RGB/PBR asset plus normal and RGB turntables | How good, plausible, and image-aligned is the asset when no unique GT exists? |
| Hi3DBench image | 510 official image prompts when the gated archive is available | Official view layout and reference image | How do object, part, material, coherence, and alignment quality compare? |
| T3Bench text | 300 official prompts | Native text-conditioned asset | How do quality and alignment change across prompt complexity? |
| Hi3DBench text | 510 official text prompts | Native text-conditioned asset and official views | How does hierarchical text-to-3D quality compare? |
| Production health | Every attempted case in every track | Artifact or explicit failure row | Is the model practical, reproducible, and export-safe? |
| Scorecard | Metrics | Valid tracks | What it cannot claim |
| Reference geometry | MD + MD-F1; CD + CD-F1; normal cosine; PSNR-N; LPIPS-N | Paired-GT only | Open-domain plausibility or human preference |
| Paper alignment | CLIP; CLIP-N; ULIP-2; Uni3D; optional OpenShape | Open image and paired-GT | Absolute geometric correctness |
| Rendered appearance | CLIP/DINO FD and KD; LPIPS/PSNR where paired views exist | Track-dependent | Mesh validity or hidden-surface accuracy |
| Human-aligned quality | Hi3DEval hierarchy; 3DGen-Score five dimensions; human preference subset | Official RGB/normal view contracts | Ground-truth point/mesh distance |
| Text evaluation | T3Bench quality and alignment; Hi3DEval text hierarchy | Native text-to-3D only | Native capability for image-bridged systems |
| Production health | Success rate; finite vertices; components; watertightness; texture coverage; latency; peak VRAM | Every attempted case | Visual or semantic quality |
| Paper | Final-generation evaluation | Paired reconstruction / geometry evaluation | Suite decision |
| TRELLIS.2 | 100 AI images; CLIP, CLIP-N, ULIP-2, Uni3D, RGB/normal preference | Toys4K + 90 curated assets; MD/CD with F-scores and normal PSNR/LPIPS | Use both ideas as separate tracks; do not call the public open set an exact reproduction |
| TRELLIS-1 | Toys4K generation distribution metrics, CLIP, and user study | 500 Toys4K assets; CD, F-score, PSNR-N, LPIPS-N | Adopt its reference geometry definitions and fixed-view discipline |
| Direct3D-S2 | ULIP-2, Uni3D, OpenShape, and a 75-mesh user study | VAE reconstruction is primarily qualitative | Keep as protocol evidence; optional runner |
| Hunyuan3D-2.1 | ULIP-I/T and Uni3D-I/T for shape; FID, CLIP-FID, LPIPS for texture | No final-shape MD/CD table | Evaluate its final mesh with the suite's external paired-GT metrics |
| Layer | Owner | Integration contract |
| Inference and artifact production | WilD3DGen suite | Immutable case IDs, model revision, seed, mesh/GLB, RGB/normal sidecars |
| Reference geometry | WilD3DGen suite | Normalized GT mesh, camera registration, MD/CD/F-score/normal definitions |
| Hierarchical quality | Official Hi3DEval adapter | Preserve checkpoint, prompt modality, view layout, and per-dimension outputs |
| Human-preference prediction | Versioned 3DGen-Score adapter | Preserve checkpoint revision, preprocessing, pairwise/absolute semantics, and five dimensions |
| Text complexity | Official T3Bench adapter | Native text prompts and official quality/alignment aggregation |
| Track / component | Current evidence | Canonical interpretation | Open gate |
| Toys4K runs | Five model families have 3,229 records; TRELLIS.2 was partially complete at the 2026-08-11 audit | Generation records exist, but old mixed proxy metrics are not final leaderboard results | Re-score only after the paired-GT protocol is frozen |
| Former TRELLIS.2 proxy | TRELLIS.2 and Hunyuan3D-2.1 each have 476 generated cases | Rename to paired-GT pilot; values are not comparable to TRELLIS.2 Table 2 | Add MD, MD-F1, CD-F1, PSNR-N, LPIPS-N and canonical alignment |
| T3Bench | 300/300 native TRELLIS-1 records; bridge experiments also exist | Native and bridged systems must remain separate | Complete official quality and alignment adapters |
| Hi3DBench text | 510/510 generation records for available systems | Generation complete does not imply metric complete | Complete the official hierarchy and equalize coverage |
| Hi3DBench image | Official 510-image archive unavailable locally | Blocked input track, not a zero score | Obtain the official archive and verify its license/provenance |
| 3DGen-Score | No accepted local adapter yet | Optional human-aligned metric family | Reproduce official calibration examples in an isolated Miniconda environment |
| Gate | Deliverable | Exit criterion |
| 1 · Freeze data | Versioned manifests, licenses, IDs, source images, prompts, GT assets, train-overlap audit | Every model consumes the same immutable case list |
| 2 · Freeze geometry | Normalization/alignment, outer-shell extraction, MD/CD/F-score thresholds, point counts, cameras | Two independent implementations agree on the calibration set |
| 3 · Freeze artifacts | Mesh/GLB, PBR channels, RGB/normal view layouts, metadata and explicit failure schema | All model adapters pass schema validation |
| 4 · Integrate official metrics | CLIP/CLIP-N, ULIP-2, Uni3D, LPIPS, Hi3DEval, T3Bench, optional 3DGen-Score | Checkpoint and preprocessing provenance are emitted per row |
| 5 · Calibrate | Small fixed cross-model subset with render grids and hand-checked geometry | No metric inversion, camera mismatch, or silent missing-data averaging |
| 6 · Full execution | All applicable models × all full manifests with resume and failure accounting | Case-level JSONL, aggregates, confidence intervals, and galleries exist |
| 7 · Publish | Separate geometry, alignment, human-quality, text, and operations scorecards | Every table declares N, failures, N/A, metric version, model revision, and runtime hardware |