← Back to project

WilD3DGen · Canonical Complete 3D Benchmark Suite

Merged suite specification, metric provenance, repository integration, and implementation gates · updated 2026-08-11
benchmarkimage-to-3Dtext-to-3Dpaired GTHi3DEval3DGen-ScoreT3Bench

TL;DR

1 · Canonical suite topology

Canonical full-suite topology Final model outputs are evaluated once, then routed to independent scorecards. PAIRED-GT IMAGE → 3D Toys4K · fixed 3,229 pool final generated mesh vs source GT OPEN IMAGE → 3D TRELLIS.2-style 100 · Hi3DBench 510 no unique GT geometry NATIVE TEXT → 3D T3Bench 300 · Hi3DBench 510 native text checkpoints only ONE VERSIONED ARTIFACT CONTRACT mesh / GLB · RGB turntable normal turntable · exact prompt/image model revision · seed · camera latency · peak VRAM · failure record unsupported modalities remain N/A REFERENCE GEOMETRY MD/CD/F-score · PSNR-N/LPIPS-N PAPER ALIGNMENT CLIP/CLIP-N · ULIP-2 · Uni3D HUMAN-ALIGNED QUALITY Hi3DEval · 3DGen-Score · preference PRODUCTION HEALTH validity · failures · latency · VRAM No universal score: publish separate scorecards with confidence intervals and missing-data counts.

2 · Full benchmark tracks

TrackFull splitRequired outputPrimary question
Paired-GT final image-to-3DFixed Toys4K pool: 3,229 public assets; deterministic source viewsFinal generated mesh/GLBHow accurately does the complete model recover known 3D geometry from one image?
Open-image final generation100 fixed public/licensed AI images in the TRELLIS.2 styleFinal RGB/PBR asset plus normal and RGB turntablesHow good, plausible, and image-aligned is the asset when no unique GT exists?
Hi3DBench image510 official image prompts when the gated archive is availableOfficial view layout and reference imageHow do object, part, material, coherence, and alignment quality compare?
T3Bench text300 official promptsNative text-conditioned assetHow do quality and alignment change across prompt complexity?
Hi3DBench text510 official text promptsNative text-conditioned asset and official viewsHow does hierarchical text-to-3D quality compare?
Production healthEvery attempted case in every trackArtifact or explicit failure rowIs the model practical, reproducible, and export-safe?

3 · Model and modality matrix

ModelImage → 3DNative text → 3DSuite role
EVA01YESCheckpoint-dependentCurrent comparison model
UniLat3DYESN/ACurrent comparison model
Pixal3DYESN/ACurrent comparison model
TRELLIS-1YESYESCurrent comparison model + protocol source
TRELLIS.2YESN/ACurrent comparison model + protocol source
Hunyuan3D-2.1YESN/ACurrent comparison model + protocol source
WilD3DGenYESReleased-checkpoint-dependentTarget model
Direct3D-S2YESN/AOptional reference baseline; official-paper protocol source

4 · Metric coverage and reporting policy

ScorecardMetricsValid tracksWhat it cannot claim
Reference geometryMD + MD-F1; CD + CD-F1; normal cosine; PSNR-N; LPIPS-NPaired-GT onlyOpen-domain plausibility or human preference
Paper alignmentCLIP; CLIP-N; ULIP-2; Uni3D; optional OpenShapeOpen image and paired-GTAbsolute geometric correctness
Rendered appearanceCLIP/DINO FD and KD; LPIPS/PSNR where paired views existTrack-dependentMesh validity or hidden-surface accuracy
Human-aligned qualityHi3DEval hierarchy; 3DGen-Score five dimensions; human preference subsetOfficial RGB/normal view contractsGround-truth point/mesh distance
Text evaluationT3Bench quality and alignment; Hi3DEval text hierarchyNative text-to-3D onlyNative capability for image-bridged systems
Production healthSuccess rate; finite vertices; components; watertightness; texture coverage; latency; peak VRAMEvery attempted caseVisual or semantic quality
The suite publishes independent physical, perceptual, preference, and operational scorecards so one strong family cannot conceal another failure mode.

5 · What the model papers actually evaluate

PaperFinal-generation evaluationPaired reconstruction / geometry evaluationSuite decision
TRELLIS.2100 AI images; CLIP, CLIP-N, ULIP-2, Uni3D, RGB/normal preferenceToys4K + 90 curated assets; MD/CD with F-scores and normal PSNR/LPIPSUse both ideas as separate tracks; do not call the public open set an exact reproduction
TRELLIS-1Toys4K generation distribution metrics, CLIP, and user study500 Toys4K assets; CD, F-score, PSNR-N, LPIPS-NAdopt its reference geometry definitions and fixed-view discipline
Direct3D-S2ULIP-2, Uni3D, OpenShape, and a 75-mesh user studyVAE reconstruction is primarily qualitativeKeep as protocol evidence; optional runner
Hunyuan3D-2.1ULIP-I/T and Uni3D-I/T for shape; FID, CLIP-FID, LPIPS for textureNo final-shape MD/CD tableEvaluate its final mesh with the suite's external paired-GT metrics

6 · Combined repository strategy

LayerOwnerIntegration contract
Inference and artifact productionWilD3DGen suiteImmutable case IDs, model revision, seed, mesh/GLB, RGB/normal sidecars
Reference geometryWilD3DGen suiteNormalized GT mesh, camera registration, MD/CD/F-score/normal definitions
Hierarchical qualityOfficial Hi3DEval adapterPreserve checkpoint, prompt modality, view layout, and per-dimension outputs
Human-preference predictionVersioned 3DGen-Score adapterPreserve checkpoint revision, preprocessing, pairwise/absolute semantics, and five dimensions
Text complexityOfficial T3Bench adapterNative text prompts and official quality/alignment aggregation
The execution harness and preference evaluator are complementary layers; the artifact contract is their stable integration boundary.
The local suite is operationally broader, while 3DGen-Bench adds explicit preference supervision; neither replaces paired-GT geometry.
A single immutable generation output fans out to official paper metrics, reference geometry, human-aligned scorers, and operational auditing.

7 · Current state, corrected labels

Track / componentCurrent evidenceCanonical interpretationOpen gate
Toys4K runsFive model families have 3,229 records; TRELLIS.2 was partially complete at the 2026-08-11 auditGeneration records exist, but old mixed proxy metrics are not final leaderboard resultsRe-score only after the paired-GT protocol is frozen
Former TRELLIS.2 proxyTRELLIS.2 and Hunyuan3D-2.1 each have 476 generated casesRename to paired-GT pilot; values are not comparable to TRELLIS.2 Table 2Add MD, MD-F1, CD-F1, PSNR-N, LPIPS-N and canonical alignment
T3Bench300/300 native TRELLIS-1 records; bridge experiments also existNative and bridged systems must remain separateComplete official quality and alignment adapters
Hi3DBench text510/510 generation records for available systemsGeneration complete does not imply metric completeComplete the official hierarchy and equalize coverage
Hi3DBench imageOfficial 510-image archive unavailable locallyBlocked input track, not a zero scoreObtain the official archive and verify its license/provenance
3DGen-ScoreNo accepted local adapter yetOptional human-aligned metric familyReproduce official calibration examples in an isolated Miniconda environment
Existing generation coverage is useful input, but metric completeness and protocol correctness remain separate release gates.

8 · Full-version implementation gates

GateDeliverableExit criterion
1 · Freeze dataVersioned manifests, licenses, IDs, source images, prompts, GT assets, train-overlap auditEvery model consumes the same immutable case list
2 · Freeze geometryNormalization/alignment, outer-shell extraction, MD/CD/F-score thresholds, point counts, camerasTwo independent implementations agree on the calibration set
3 · Freeze artifactsMesh/GLB, PBR channels, RGB/normal view layouts, metadata and explicit failure schemaAll model adapters pass schema validation
4 · Integrate official metricsCLIP/CLIP-N, ULIP-2, Uni3D, LPIPS, Hi3DEval, T3Bench, optional 3DGen-ScoreCheckpoint and preprocessing provenance are emitted per row
5 · CalibrateSmall fixed cross-model subset with render grids and hand-checked geometryNo metric inversion, camera mismatch, or silent missing-data averaging
6 · Full executionAll applicable models × all full manifests with resume and failure accountingCase-level JSONL, aggregates, confidence intervals, and galleries exist
7 · PublishSeparate geometry, alignment, human-quality, text, and operations scorecardsEvery table declares N, failures, N/A, metric version, model revision, and runtime hardware

9 · Primary sources

Canonical plan · merged from the complete-suite and repository-comparison reports · no fast-eval leaderboard · updated 2026-08-11