← all projects

3D Generation & Editing

GitHub: weikaih04/3d-edit-engine · GitHub: weikaih04/vlm-quality-complexity-judge
31 item(s)
REPORT
Stage A · 3D viewer (real meshes)
orbit-able source/GT/ours meshes, camera-linked; geometry frozen so only the PBR differs
2026-07-30 · static HTML
REPORT
Stage A · instruction-based 3D editing with a source-asset rail
0/6 to 6/6 instruction following; the zero-init trap that silently disconnected the source rail; detail preservation still open
2026-07-30 · static HTML
REPORT
Why do our generations look oily? A four-layer material audit
Demo outputs read as chrome (GLB metallic 0.996). Audited data/model/decode/export: all four clean (GT 0.016, export 0.002, tex flow 0.000, e2e 0.004). In-distribution 8-asset fidelity: pred 0.078 vs GT 0.016 (1/8 outlier). REFUTED the specular-highlight hypothesis: no photometric perturbation moves metallic. Failure is OOD on real photos (0.49-0.996). Proposes VLM material-semantics conditioning (Seed3D 2.0 validates the route; our caption data is already GT-PBR-grounded).
2026-07-30 · static HTML
APP :7860
Tri-modal 3D demo (S3 40k) — image / multi-view / text
Live demo of the S3 tri-modal 40k model: upload an image, 2-4 views, or type a description, get a textured GLB. Conditioning is computed live (v2.2 VLM + DINOv3), not read from a cond cache. ~20-25s/request on one H200; cold start ~2 min. Launcher srun's into a held node and opens a Cloudflare tunnel (URL in runs/demo_url.txt).
2026-07-28 · served app
cd /fsx/home/weikai.huang/3dgen/model/BLIP3o && bash scripts/run_demo_app.sh
open :7860
REPORT
Tri-modal viewer: one model from image / 4 views / text
Interactive per-asset comparison of the S3 unified model generating the same held-out object from a single image, four views, and a text caption alone. Full cascade, identical seed, at 40,000 cumulative steps (SS/shape/texture all ckpt-37000 resumed from 3k). Mean SS-IoU 0.407 / 0.388 / 0.205 (0.361 / 0.341 / 0.143 excluding a memorised duplicate). Four views still lose to one image on 11/14 assets, but they now beat a 4-copy control by +0.020 IoU -- a real effect that has been flat since ~21k steps. Each of the I1/IM/T columns flips between the render and the actual generated mesh, orbit-able and camera-linked.
2026-07-27 · static HTML
REPORT
WilD3DGen Complete 3D Benchmark Suite Plan
Visual full-version evaluation plan: suite topology, model modality matrix, metric coverage, Hi3DEval/T3Bench integration, and implementation gates.
2026-07-25 · static HTML
REPORT
Part-level co-move groups: geometry containment + VLM semantic merging
Two-layer part-quality stack: geometry (winding containment / nested-shell clustering / tangent-continuity) handles hidden interior & double walls; VLM (color-code clusters -> one Qwen call) groups+names semantic parts. Fixes 'eyes left behind moving head'. E3 remove 9->12/12, E2 semantic-add 93%.
2026-07-22 · static HTML
REPORT
Multi-image (4-view) conditioning for 3D — the complete investigation
Full arc: info exists (+28pts visibility) but SS ignores it (4d≈4c). 6 nulls (IM probe, view-randomization, TRELLIS-1 aggregation, VGGT-REPA v1/v2/v3). The one win = scaled sincos view-embed: flips attention from content-blind to visibility-aligned routing (top1 0.53→0.65) but SATURATES at +0.022 contrast with no absolute gain (4v vs 1v +0.001, still <I1 0.404). Root cause: 16^3 occupancy objective never rewards fusion. Mechanism != incentive.
2026-07-22 · static HTML
REPORT
500k Synthetic 3D Pipeline — Pixal3D (interactive, merged)
Merged interactive viewer: (1) end-to-end products (Qwen gen → rembg → QC 99/300 → image→3D 99/99), (2) Pixal3D on 10 everyday LVIS concepts, (3) vanilla TRELLIS.2 vs Pixal3D on 6 coverage-gap concepts. Tabbed, drag-rotate GLBs.
2026-07-22 · static HTML
REPORT
Text-to-3D Coverage · Gate-1 Pilot — T2I→3D synthesis works
Z-Image vs Qwen-Image → TRELLIS.2 image→3D on 10 coverage-gap concepts. Verdict: PASS — no synth domain gap, thin topology reconstructs (gyroid/knots), bottleneck moved to T2I semantics → VLM QC gate mandated; wireframe-cavity = the one failure.
2026-07-19 · static HTML
REPORT
Text-to-3D feasibility — WilD3DGen
Text-only conditioning of the 3-stage flow cascade: architecture validity, 3 OOD rounds, convergence (LR-reset test), n=58 quantitative eval (text CLIP 0.239 vs GT 0.248), and the T2I-bridge deployment path.
2026-07-14 · static HTML
REPORT
3d-gen · Why VLM hiddens fail as spatial conditioning — diagnosis, fixes, and internalization
RoPE position-blindness probes; 6-arm conditioning ablation (crossdadd 0.380 ≈ fusion at half cond length); prod DinoPosStamp +25% (0.256→0.319); DINO-distilled VLM internalization with rollout dissociation; KD negative result
2026-07-13 · static HTML
REPORT
Material-Import Bug Audit (renders_cond)
91,814 assets (8.4%) have grey renders but colorful PBR materials. Detected via cond_sat vs decoded-PBR gt_sat. Concentrated in TexVerse(11.3%)+SketchfabV1(9.7%). GT recovered from latents via shape_dec->tex_dec octree transfer.
2026-07-09 · static HTML
REPORT
Editing Data Engine — 12-task pilot pairs (interactive QC viewer)
All 12 edit types instantiated at pilot scale (5-10 pairs each): per-pair instruction, claude verdict, 2D QC strip, drag-to-rotate before/after GLBs (draco). v0 — E1a good-view redo + E1c/E5 land with batch-2.
2026-07-10 · static HTML
REPORT
Editing Data Engine Plan: 440K curated assets → ~750K instruction-editing pairs
Design doc: 10-type edit taxonomy mapped to prior-work pipelines (Nano3D/EVA01/Steer3D/UniVerse3D), two-track engine, budget, pilot gate
2026-07-08 · static HTML
REPORT
Interactive 3D viewer — good-view fusion on held-out (drag/rotate GLB)
6 held-out assets, GT vs generated (full flow chain), PBR GLBs in model-viewer.
2026-07-08 · static HTML
REPORT
VLM Quality+Complexity Judge (ready_v4)
Qwen3.6-27B judged 1.04M PBR assets: 6 axes (3 quality+3 complexity) + style + issues. yes=440k near 500k target. Interactive turntable gallery.
2026-07-06 · static HTML
REPORT
Model-arch ideas: unified geometry+texture SLAT diffusion
Design study: light-unify (frozen VAEs, 64ch joint flow, ~2-3 days) vs full Uni-VAE (~3 wks); PBR-specific risk of naive joint denoising; staggered noise schedules with independent-t training make delta an inference-time cascade↔joint knob.
2026-07-06 · static HTML
REPORT
Shape-512 conditioning ablation (cross-attn / MMDiT / per-layer KV)
3-arm ablation: read-once cross-attn baseline wins (0.302); MMDiT co-evolving dual-stream marginally worst (0.327), deficit is architectural not init (fairness-corrected). Aligns with DiT-Air.
2026-07-05 · static HTML
REPORT
Diagnosing & Fixing Mode Collapse — discrete-token 3D VLM
The debugging journey: image->3D collapsed to boxes because 5824 tokens let the model ignore the image (image-usage GAP flat at 0.04). Fix = short codes 1536 + input-noising + sampling + light CFG. GAP 0.04->0.12, faithful coarse generation, VQA 92% preserved. Includes GAP curve + before/after renders.
· static HTML
REPORT
Asset Gallery — v2/v3/v4 sampled 3D assets (multi-view)
Interactive gallery of 930 quality-sampled assets across all subsets (github/sketchfab/TexVerse/ABO/HSSD/Toys4k + v4 Sketchfab-v1). Each tile = 4 views; filter by subset; click to enlarge. Missing-texture(magenta)+degenerate renders filtered out.
2026-07-01 · static HTML
REPORT
VLM-3D Stage-1 · Method, Data & Results
Full system: 3-component frozen VQ tokenizer (5824 tok/asset) + Qwen3.5-2B full-backbone AR + 1.35M data mixture (3D 65% / replay 35%) + examples + 1-epoch eval
2026-06-30 · static HTML
REPORT
Ready Training Data — v2 / v3 / v4: summary & examples
Counts by modality for ready_v2/v3/v4 + per-subset rendered examples + first-hand quality feedback (texture-coverage gap, dark/small renders, TexVerse mesh hygiene, v4 character-skew). v4=Sketchfab-v1 426,521 selected, pre-Phase-2.
2026-06-30 · static HTML
REPORT
VLM-3D Stage-1 · ~27× speedup, 31.3% MFU (fla + packing)
Full arc: ~39 s/it (naive) -> 5.7 (FA3+bs4) -> 1.43 (fla+packing); MFU 1.2%->8.6%->31.3% on 8xH200
2026-06-30 · static HTML
REPORT
SLAT flow training speed — profiling & optimization
LOSSLESS true MFU ~10%->~30%, ~3x throughput: broadcast fix + Triton fusion + torch.compile(SS) + FA3 + bs8
2026-06-29 · static HTML
REPORT
Caption v2 — holistic + per-part PBR (Qwen3.6)
All captions regenerated with Qwen3.6-27B: holistic 453,977 + new per-part PBR texture 296,387. OLD vs NEW samples.
2026-06-29 · static HTML
REPORT
3D Tokenizer Quality: SS + SLAT, and the codes-vs-cells law
All three discrete tokenizers (SS structure + shape/texture SLAT) on frozen TRELLIS.2. Key finding: codes-vs-cells law — SS (high-freq boundary) wants more spatial cells (12³=1728 tok, occ-IoU@32 0.88); SLAT (smooth features) wants more codes/cell (8³×ng4 ConvNeXt, 2048 tok). Measure SS at 32³ not 64³.
2026-06-25 · static HTML
REPORT
MolmoAct2 -> BLIP3o-NEXT: VLM-side 3D pretraining (VP1+VP2) + per-layer KV
Design/hypothesis: two new VLM-side stages (3D-understanding + editing pretrain) + per-layer KV conditioning, adapting MolmoAct2. Not yet validated.
2026-06-20 · static HTML
REPORT
3D Data Overview — June 2026 milestone
4 datasets (v1 Sketchfab 12 TB, v1 GitHub 1.9 TB, v2 SWH 11.1 TB, TRELLIS.2 18.6 TB) + 30 TB R2 backup verified.
2026-06-17 · static HTML
REPORT
VLM-Native 3D Conditioning (exploration summary)
Project arc: condition swap, infra, DINOxQwen fusion, multi-task.
2026-06-17 · static HTML
REPORT
S3 tri-modal unified training: config, data and the four bugs found in setup
Config of record for the I1+IM+T unified run on 400k: rebuilt IM cache (weighted 4-of-16 views), 1.69M caption entries, xf2 connector, sincos view embed, pos_stamp off, view-count randomization restored, new qwen dropout. No MDS needed. Four setup bugs documented.
2026-07-25 · static HTML
Add a report to this project (any agent with the publish-milestone-viz skill):
1. git clone https://github.com/weikaih04/research-proj-report.git
2. render a self-contained HTML (see render_report.py)
3. python hub.py add-report --project 3d-generation-editing --title "..." --file your.html
4. git add -A && git commit -m "publish" && git push → Cloudflare Pages redeploys.