← Back to project

WilD3DGen · Complete 3D Benchmark Suite

Evaluation-only plan · full-version protocol · 2026-07-25
benchmarkimage-to-3Dtext-to-3DHi3DEvalvisual plan

TL;DR

1 · Suite topology

Complete benchmark topologyEvery model produces the same artifact schema; unsupported modalities stay N/A.IMAGE INPUTToys4K · 3,218Hi3DBench image · 510TEXT INPUTT3Bench · 300Hi3DBench text · 510GT GEOMETRYHard-OOD · 100Objaverse / Sketchfab7-MODEL COMPARISONEVA01 · UniLat3D · Pixal3DTRELLIS-1 · TRELLIS.2Hunyuan3D-2.1WilD3DGenIMAGE SCORECARDCLIP · DINO · ULIP/Uni3DLPIPS · novel viewsTEXT SCORECARDT3Bench · Hi3DEvalalignment · qualityGEOMETRY + HEALTHChamfer · F-score · normalsvalidity · VRAM · latency

2 · Complete suite

TrackFull splitPurposeModels
Toys4K image3,218Common single-view image-to-3D comparisonAll seven
Hi3DBench image510 image promptsObject, part, material, and spatial qualityAll seven
Hard-OOD geometry100 fixed GT assetsChamfer, F-score, normals, and novel-view reconstructionAll seven
T3Bench text300 promptsSingle object, surroundings, and multiple objectsNative text models
Hi3DBench text510 text promptsHierarchical text-to-3D qualityNative text models
Production healthEvery artifactExport validity, latency, VRAM, and failure rateEvery model/track

3 · Model × modality matrix

ModelImage-to-3DNative text-to-3DPrimary evaluation
EVA01YESYESImage + text tracks
UniLat3DYESN/AImage tracks
Pixal3DYESN/ASingle-view common track
TRELLIS-1YESYESImage + text tracks
TRELLIS.2YESN/AImage tracks
Hunyuan3D-2.1YESN/AImage tracks
WilD3DGenYESAccording to released checkpointTarget model

4 · Metric coverage

TrackImplemented foundationAdapter required
Toys4K imagePSNR/MAE, SSIM optional, mesh health, JSONL aggregationCLIP, DINOv2 FD/KD, ULIP-2, Uni3D, LPIPS
Hi3DBench image/textArtifact protocol and aggregationOfficial Hi3DEval object/part/material evaluator
Hard-OOD geometryChamfer-L2, artifact loadingF-score, normal consistency, fixed GT sampling
T3Bench textArtifact protocol and aggregationOfficial T3Bench quality/alignment evaluator
Production healthMesh existence, finite vertices, watertightness, componentsTexture coverage, mesh size, peak VRAM, failure taxonomy
FOUNDATION READYOFFICIAL ADAPTERSFULL RUN

5 · Full implementation gates

PhaseDeliverableExit criterion
1Freeze manifests, prompt IDs, cameras, seeds, and GLB schemaEvery model consumes identical manifests
2Finish deterministic geometry/render/health metricsClean-environment run succeeds
3Integrate CLIP, DINOv2, LPIPS, ULIP-2, and Uni3DCheckpoint and preprocessing recorded
4Integrate official Hi3DEval and T3BenchNative image/text tracks score successfully
5Run all seven models on complete manifestsJSONL, aggregates, failures, and render grids exist
6Publish scorecards and qualitative galleriesNo silent missing-data averaging

6 · Sources and decisions

Plan status: registry updated to complete_version_only; learned-metric and official-evaluator adapters remain implementation gates.