← Back to project

WildDet3D · NeurIPS 2026 Rebuttal Tracker

AC-level response strategy · completed experiments · reviewer-by-reviewer evidence map · 2026-07-27
rebuttalNeurIPS 2026experiment planningevidence audit

1 · Decision context

Current decision risk. The meta-review is leaning reject because it frames the work as primarily semi-automatic dataset construction and questions whether the model-assisted labels are independent of MoGe-2. Limited architectural complexity is secondary: the AC explicitly says a simple architecture is acceptable when the evidence is convincing.
AC-level position. We respectfully disagree with that framing. WildDet3D studies one general-purpose detector that jointly supports 13.5K in-the-wild categories, text/point/box/exemplar interaction, and optional partial or full depth without retraining. WildDet3D-Data is enabling supervision for this setting, not the sole contribution.
Rebuttal pillarCore claimBest evidence
1 · In-the-wild scaleA unified 3D detector over open-ended categories and four prompt interfaces, rather than a small closed-vocabulary benchmark. 13.5K categories; 230K human-verified boxes; completed four-prompt evaluation.
2 · Optional depthThe algorithmic insight is an RGB-first model whose same weights can opportunistically use partial/full metric depth. 70/20/10 depth dropout, zero-initialized residual fusion, Stereo4D/DROID gains, and the new MoGe-2→UniDepthV2 hidden-GT counterfactual.
3 · Downstream utilityThe interface transfers to real stereo and embodied settings, not only the paper's derived benchmark. DROID 6.8→24.1 zero-shot with depth and 26.9→32.0 after adaptation; iPhone, Quest, Franka, and grounded-VLM deployments.

Writing stance: push back firmly but professionally. Agree that monocular metric scale is physically underdetermined; clarify that this is precisely why optional sensor conditioning is central. Do not claim that RGB eliminates scale ambiguity.

Review record · Verbatim original

Verbatim review snapshot. The text below is reproduced from the author-provided OpenReview export. Original wording, spelling, scores, and timestamps are preserved. Our response plans appear in separate sections.
Open full AC meta-review and official reviewer record
Meta Review of Submission10926 by Area Chair PPk5
Meta Reviewby Area Chair PPk519 Jul 2026, 11:55 (modified: 23 Jul 2026, 10:42)Senior Area Chairs, Area Chairs, Authors, Reviewers Submitted, Program Chairs, Area Chair PPk5[Revisions](https://openreview.net/revisions?id=N8y5GX0R0O)
Metareview:
This paper proposes an architecture that detects objects in 3D (more exactly predicts a 3D bounding box) from images and a dataset, which is used to train this architecture. The architecture takes various types of prompts (similar to the ones for SAM) and is relatively standard now. The dataset is large and created semi-automatically from existing 2d datasets (COCO, LVIS, Obj365, V3Det).

Strength: The model and the dataset are potentially very valuable.

Weaknesses: The reviewers are not convinced that the dataset annotations are really high quality. Rev bVkw complains that the technical novelty is limited--this is not a problem for the AC (no need to invent a fancy architecture if a simpler one is just fine) however, the contribution is then more on the data creation side, which is not really suitable to NeurIPS.

Because of these weaknesses, the paper is aiming toward rejection at this stage.

Add:
Official Review of Submission10926 by Reviewer Z8Kd
Official Reviewby Reviewer Z8Kd28 Jun 2026, 08:14 (modified: 23 Jul 2026, 07:23)Program Chairs, Senior Area Chairs, Area Chairs, Reviewers Submitted, Authors, Reviewer Z8Kd[Revisions](https://openreview.net/revisions?id=FHNaJxfA81)
Summary:
This paper presents WildDet3D, an open-world monocular 3D object detector with optional depth as input. WildDet3D is designed to support different types of prompts, including text, box, point, and exemplar prompts, in one architecture. The architecture includes a RGB image encoder, a RGB-D encoder, a depth fusion module with zero-conv, followed by 2D/3D detection heads. Crucially, the authors propose a new large-scale dataset, WildDet3D-Data, to train the WildDet3D detector. WildDet3D-Data is built by lifting existing 2D annotations from various datasets (COCO, LVIS, Obj365, V3Det) into 3D using multiple existing 3D detection methods, and then selecting/filtering the candidates through human annotation and VLM scoring. The final dataset covers many more categories than existing 3D detection datasets. The paper reports strong results on WildDet3D-Bench, Omni3D, zero-shot AV2/ScanNet, and two additional benchmarks for stereo and embodied settings.

Contribution Type: General: Most submissions will fall into this type.
Strengths And Weaknesses:
Strengths
The paper addresses an important problem of open-world 3D detection from single images, which is useful for robotics, AR/VR, and embodied AI. Supporting different types of prompts and in one model is a practical direction.
The effort of making the WildDet3D-Data dataset is very impressive. The authors propose a human-in-the-loop pipeline to select the candidate annotations, and the final dataset delivers a large category coverage. The ablation studies show that the dataset can be potentially very useful to the community.
Weaknesses
The promptable claim is broader than the evaluation shows. The paper claims support for text, box, point, and exemplar prompts, but the main quantitative evaluation only covers text and box prompts. I also did not see the qualitative results of point and exemplar prompts in the paper or from the supplement videos. Although the authors may not be able to compare against other methods on point and exemplar prompts, showing quantiative numbers as an ablation study to compare with box and text prompts as well as qualitative results are crucial to support the claim of the paper.
I am not convinced that the WildDet3D-Data is a high quality dataset with accurate 3D bounding boxes. WildDet3D-Data seems to be built from lifted 2D annotations using candidate 3D box generators plus human/VLM selection, rather than from independent metric 3D ground truth. This could be useful, but can also be very noisy, as we can clearly see from the GT row in Figure 5, and potentially biased toward methods similar to the candidate generators. In Table 8, the authors provide a dataset ablation study and show the trend of using WildDet3D-Data for training is useful. But critically it uses WildDet3D-Bench, which is created using the same annotaion pipeline, as the evaluation ground truth, which can bias the evalution toward the WildDet3D-Data training. In the table it also shows that using "Other 3D" did not help the evaluaiton on WildDet3D-Bench, making me more suspicious that this bias is real.
Quality: 3: good
Clarity: 3: good
Significance: 2: not good
Originality: 2: not good
Questions:
Can the authors provide quantiative evaluations on the point and examplar prompts?
How independent is WildDet3D-Bench from the annotation pipeline? Since the 3D labels are selected from candidates produced by existing models and geometry heuristics, could the benchmark favor methods trained on similar generated labels? Can the authors provide a similar ablation as in Table 8, but using other more accurate datasets, such as Omni3D and/or ScanNet, as the evaluation?
Limitations:
yes

Rating: 3: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly.
Confidence: 4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Ethical Concerns: NO or VERY MINOR ethics concerns only
Paper Formatting Concerns:
none

Code Of Conduct Acknowledgement: Yes
Responsible Reviewing Acknowledgement: Yes
Add:
Official Review of Submission10926 by Reviewer qeMe
Official Reviewby Reviewer qeMe23 Jun 2026, 15:16 (modified: 23 Jul 2026, 07:23)Program Chairs, Senior Area Chairs, Area Chairs, Reviewers Submitted, Authors, Reviewer qeMe[Revisions](https://openreview.net/revisions?id=1SbXkgQ6nJ)
Summary:
WildDet3D targets open-world monocular 3D detection with two contributions: (1) a geometry-aware architecture that takes a single RGB image, accepts text/point/box/exemplar prompts through a shared encoder, and treats depth as an optional inference-time refinement and (2) WildDet3D-Data, a 1M-image open-vocabulary 3D dataset (13.5K categories, 230K human-verified boxes via 1,786 Prolific annotators plus ~900K VLM-scored), built by lifting 2D detection datasets to 3D through five candidate generators, multi-stage filtering, and a two-path human/VLM selection pipeline. They report SOTA on Omni3D (34.2/36.4 text/box, 10× fewer epochs), large zero-shot gains on AV2/ScanNet, and a new in-the-wild benchmark plus stereo and embodied benchmarks.

Contribution Type: General: Most submissions will fall into this type.
Strengths And Weaknesses:
Strengths

High Quality data: 138× Omni3D's category count, with a human-in-the-loop pipeline (gold-task QC, 84–98% pass rates) rather than VLM-only labeling albiet VLM scoring.

+4.2 on Omni3D, +16.5/+17.4 ODS zero-shot, and a near-doubling with depth are well outside noise

Weaknesses

On WildDet3D-Bench the depth signal and the ground-truth 3D boxes are both derived from the same MoGe-2 monocular depth via the same lifting pipeline. So "add depth -> 22.6 -> 41.6" partly measures agreement between the model's depth input and depth-derived labels, not independent geometric gain.

WildDet3D's gains conflate a new architecture, a new training set, and strong pretrained backbones (SAM 3 ViT-H, a metric-depth DINOv2). To know what the architecture contributes, one wants a baseline trained on the same WildDet3D-Data, or WildDet3D trained on only Omni3D vs 3D-MOOD on Omni3D at matched backbone capacity

Quality: 3: good
Clarity: 3: good
Significance: 3: good
Originality: 3: good
Questions:
Overall, the dataset contribution is a significant one. I would have liked more backbone ablations to understand the architecture better.

Limitations:
yes

Rating: 5: Accept: Technically solid paper, with high potential value on at least one sub-area of AI or moderate-to-high impact on more than one area of AI, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations.
Confidence: 4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.
Ethical Concerns: NO or VERY MINOR ethics concerns only
Paper Formatting Concerns:
na

Code Of Conduct Acknowledgement: Yes
Responsible Reviewing Acknowledgement: Yes
Add:
Official Review of Submission10926 by Reviewer bYkw
Official Reviewby Reviewer bYkw19 Jun 2026, 04:08 (modified: 23 Jul 2026, 07:23)Program Chairs, Senior Area Chairs, Area Chairs, Reviewers Submitted, Authors, Reviewer bYkw[Revisions](https://openreview.net/revisions?id=bRrwbpFfWz)
Summary:
This paper presents WildDet3D, a promptable open-vocabulary monocular 3D object detection framework that supports multiple prompt modalities, including text, 2D points, and 2D boxes, while optionally leveraging depth cues at inference time. The authors further introduce WildDet3D-Data, a large-scale in-the-wild 3D detection dataset with human-in-the-loop annotation and VLM-assisted scaling. Extensive experiments on the proposed benchmark, Omni3D, zero-shot transfer settings, and robotic/stereo scenarios show that WildDet3D achieves strong performance over prior methods, especially when depth information is available.

Contribution Type: General: Most submissions will fall into this type.
Strengths And Weaknesses:
clear writing and figures.
extensive experiments and achive state-of-the-art performance.
introduce WildDet3D, a geometry-aware architecture that unifies text, point, box, and exemplar prompts.
introduce WildDet3D-Data, the first truly open-vocabulary 3D detection dataset—built by applying multiple complementary methods to generate candidate 3D boxes for 2D annotations from diverse detection datasets.
Quality: 3: good
Clarity: 3: good
Significance: 2: not good
Originality: 2: not good
Questions:
I personally disagree with the claim at Line 75. It is impossible to infer metric knowledge from a single RGB. Only supply depth/3D information can get metric knowledge.
It uses many VLM/3D foundation models to annotate the dataset. However, those large models are not perfect. For example, using MoGe-2 to infer metric depth is not fully accurate while the pipeline are based on the MoGe-2 depth. It is very important to analysis the depth quality and how it affects the downstream annoatation pipeline. Second, it is hard to filter per-category physical dimensions using GPT-4.1-mini (or other VLMs) due to large intra-class variance. And the depth-to-width ratio/ axis proportions are hard to define robustly because they depend on the object pose: after rotation, an object’s intrinsic height, width, or depth may align with different coordinate axes.
The technique novelty is a bit limited. The whole pipeline is built upon existing modules, such as controlnet-style module, SAM 3, DINOv2, LingBot Depth. To me, it more looks like an engineer work. Moreover, the claims at Line 74-79 remain questioned since it is hard to determin whether the performance comes from the using multiple foudation modules or their depth design.
Limitations:
The authors have discussed limitation in the supplmentary.

Rating: 3: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly.
Confidence: 5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully.
Ethical Concerns: NO or VERY MINOR ethics concerns only
Paper Formatting Concerns:
Not avaliable.

Code Of Conduct Acknowledgement: Yes
Responsible Reviewing Acknowledgement: Yes

2 · Live execution status

Main pipeline metric: loose center-distance recall. The complete five-generator annotation pipeline was replayed on 300 stratified Omni3D images containing 1,810 eligible GT objects. A prediction is successful when its 3D center is within an absolute metric threshold or a fraction of the GT cuboid body diagonal. Missing candidates count as failures. This is the primary rebuttal view requested for practical localization quality.
Pipeline viewCoverageMedian center error ≤ 1 m≤ 2 m≤ 4 m≤ 0.5× GT diag. ≤ 1× GT diag.
Raw best-of-five oracle86.0%0.36 m 69.3%77.3%82.3%70.5%80.8%
VLM selection on raw geometry85.1%0.61 m 55.3%66.7%73.6%51.5%67.0%
VLM selection after depth alignment85.5%0.88 m 45.2%50.6%52.5%31.0%46.2%

Interpretation: candidate generation has useful center localization and the VLM recovers much of the oracle envelope, but the current depth-alignment stage hurts the aggregate result—especially on outdoor images whose monocular depth scale is unstable. “VLM selection on raw geometry” is a GT-free stage-isolation counterfactual: it transfers the deployed VLM score to the same generator's pre-alignment cuboid. It is not the deployed aligned pipeline.

Supplementary strict metric. Fully oriented vis4d_cuda_ops.iou_box3d is retained as a diagnostic rather than the headline: missing candidates are zero. VLM-selected raw geometry has 14.1% coverage-adjusted mean IoU, 24.9% at IoU ≥ .25, and 8.6% at IoU ≥ .50; its same-subset best-of-five oracle is 21.8%, 39.6%, and 16.7%. After depth alignment the VLM-selected figures fall to 5.9%, 9.9%, and 1.6%, respectively.

GT-box lifting component audit (not the complete pipeline). Under official oriented IoU, monocular lifting obtains 0.270 mean / 0.182 median at 99.39% coverage; replacing the estimated depth with GT depth raises this to 0.318 mean / 0.296 median at 99.50% coverage. This isolates depth error while holding the 2D box prompt fixed.

DatasetMono mean IoU3DGT-depth mean IoU3DΔ
KITTI0.4250.399−0.027
nuScenes0.3570.362+0.004
SUNRGBD0.4490.526+0.077
Hypersim0.1170.179+0.062
ARKitScenes0.4500.532+0.082
Objectron0.4400.502+0.063

Interpretation boundary: this is a GT-box-prompted 3D lifting/component audit, not the complete five-candidate annotation pipeline. It shows that depth is a major error source on four indoor/object-centric datasets, but not the only source; the complete replay remains a separate experiment.

Independent benchmark ablation. The five finetuning mixtures use the same 12K-step budget. Row 1 is a 15,996-step pretrained initialization reference and is not presented as a sixth iso-step finetuning run. All 12 benchmark cells completed successfully.
RowRole / data mixtureOmni3D APAP15 AP25AP50ScanNet ODSCBaseNovel
1Pretrained reference (15,996 steps)32.6743.31 34.6013.8461.4364.0548.33
2Other 3D (12K)30.7740.20 32.7813.3753.4155.8241.35
3ITW human only (12K)30.4440.36 32.4012.7750.0753.2834.04
4ITW synthetic only (12K)31.3240.93 33.6113.7750.1553.6332.75
5ITW full (12K)30.5840.06 32.5513.3450.2453.4934.00
6Full mixture (12K)30.3539.76 32.2213.2655.8157.6146.78

Readout: the independent-GT evidence is mixed rather than uniformly positive. Among the matched 12K rows, ITW-synthetic is strongest on Omni3D (31.32 AP), while the full mixture is strongest on ScanNet transfer (55.81 ODSC, including 46.78 novel). ITW-full alone does not outperform Other-3D. The rebuttal should therefore claim dataset-dependent transfer benefits, not monotonic gains from every added data source.

Execution itemStateExact scope
Official cuboid auditCompleted locally2 variants × all six Omni3D test sets; 130K matched cuboids each; 2,000-bootstrap result artifacts collected. Beaker: 01KYDMDWMT16AG3XKC0V2PM6B7.
Independent data ablationCompleted locally12/12 benchmark cells completed successfully. Five matched 12K finetuning rows plus the initialization reference; 4 GPUs/task, high priority. Beaker: 01KYDMJSBE70DKQCAHN7K7ENFV.
Runtime validationCompleted locallyPython 3.11 / Torch 2.11 / CUDA 12.8 extension rebuilt; 8.2 GB checkpoint loaded; one real Omni3D forward + official evaluator completed.
Full annotation-pipeline replayCompleted locally300-image stratified Omni3D subset; SAM3 masks, metric depth, all five cuboid generators, alignment, candidate unification, Molmo2 selection, loose distance evaluation, and strict hidden-GT IoU are complete. All Beaker stages finalized successfully.
Four-prompt evaluationCompleted locallyText, point, box, and visual exemplar; matched prompt protocols, all six Omni3D test sets, official oriented-IoU AP. Corrected protocol audit completed successfully: 01KYEV4VR45Z3468PA6N0KD31Z.

2A · New result: depth-source robustness

Direct answer to the MoGe-2 circularity concern. We replayed the annotation pipeline on the same 300 Omni3D images and replaced MoGe-2 with UniDepthV2 ViT-L/14. The baseline camera calibration, candidate-generator outputs, and VLM selector scores were held fixed; independent Omni3D ground truth was hidden until final evaluation.
Hidden-GT metric (1,810 eligible objects)MoGe-2 UniDepthV2Change
Coverage85.47%84.97%−0.50 pp
Mean oriented IoU among covered objects6.92%12.58% +5.66 pp
Coverage-adjusted mean IoU5.92%10.69% +4.78 pp
IoU ≥ .15, all eligible objects15.75%28.95% +13.20 pp
IoU ≥ .25, all eligible objects9.89%18.78% +8.90 pp
Center error ≤ .5 m26.19%30.39% +4.20 pp
Center error ≤ 1 m45.25%49.83% +4.59 pp
Center error ≤ 2 m50.61%65.69% +15.08 pp

The 95% bootstrap intervals for mean IoU among covered objects are [6.24%, 7.58%] with MoGe-2 and [11.77%, 13.41%] with UniDepthV2. Coverage is unchanged, while both strict oriented IoU and loose center-distance quality improve substantially with the alternative estimator. Annotation quality is therefore not an artifact of agreement with MoGe-2.

Interpretation boundary. The absolute strict-IoU numbers should be reported transparently as an independent quality characterization, not as parity with sensor-derived ground truth. The robust conclusion is depth-source independence: changing only the estimator preserves coverage and improves final geometry.

Beaker provenance: UniDepth generation 01KYHEJ4W89ME99M3XAHZQPMJY; fixed-candidate alignment 01KYHEMFWC5PXCEZ9XJC6GZRFE; merge 01KYHETHJKTSB8ZFENJCDJN483; hidden-GT evaluation 01KYHEVNWF6GVQZWVPYYHTPAPA. All stages completed successfully in ai2/oe-encoder-data-noneng.

2B · Four-prompt evaluation

Corrected protocol. All rows use the official Omni3D oriented-IoU evaluator (AP averaged over IoU 0.05:0.50). Text, box, and exemplar use their matched checkpoint; point uses the point-trained checkpoint and click-consistent proposal selection. The point is the center of one target's GT 2D box but does not expose its extent. The visual exemplar is the first instance of a category and asks the model to retrieve all same-category instances.
PromptTarget contractOmni3D APAP15 AP25AP50Status
TextCategory query → all category instances 31.97Completed locally
PointPositive center click → one instance 29.0738.1830.5712.37Completed locally
BoxGT 2D box → one instance 36.2046.5838.0817.16Completed locally
Visual exemplarOne boxed exemplar → all same-category instances 36.1046.1337.9917.69Completed locally
PromptKITTInuScenesSUNRGBD HypersimARKitScenesObjectron
Point34.7527.5131.25 12.6356.6955.70
Box40.3734.9342.07 16.4367.7161.62
Visual exemplar38.3032.1642.81 16.6268.4562.45
Readout. All four interfaces are quantitatively functional. Point prompting reaches 29.07 AP without revealing object extent. Visual-exemplar prompting reaches 36.10 AP, essentially matching the GT-box prompt at 36.20 AP. The 7.13 AP point-to-box gap is not an evaluator failure: click-consistent proposal selection changes the final-checkpoint result by only +0.12 AP, while point-specific training raises it by +2.51 AP. A click identifies location but contains less geometric information than a box.
Negative-click diagnostic. Adding one deterministic negative click to the earlier final-checkpoint protocol lowered AP from 26.44 to 23.00. This is a non-comparable diagnostic, not one of the four headline rows; mention it only if reviewers ask about multi-click robustness.

Rebuttal-ready response. Thank you for identifying the missing evaluation. We now evaluate all four prompt modes with the official Omni3D 3D evaluator. Text, point, box, and visual-exemplar prompts obtain 31.97, 29.07, 36.20, and 36.10 AP, respectively. A single point is effective despite not exposing extent, while one visual exemplar nearly matches a GT-box prompt. We will add the qualitative results below and state each target contract explicitly.

Corrected-protocol provenance: Beaker experiment 01KYEV4VR45Z3468PA6N0KD31Z; point-trained 01KYEV4W9014KS8KV61T4FN382, box 01KYEV4WFGJEE5D2RYAHRM7G2T, exemplar 01KYEV4WPHVSDDF0TMFKQHFR4Y. All exited successfully.

Negative-click diagnostic provenance: Beaker experiment 01KYEDZH25K9S7PZPCRG8KK6R3, job 01KYEDZH5Z886AGP8SJ0HTQQ9C, same checkpoint/evaluator, 4 GPUs, exit 0.

Point-prompt 3D detection example
Point prompt. Cyan: positive click; green: prediction; white: GT.
Second point-prompt 3D detection example
Point prompt. A click localizes the requested van without exposing a box.
Visual-exemplar prompt retrieves multiple cyclists
Visual exemplar. Yellow: exemplar; green: same-category predictions.
Second visual-exemplar prompt example
Visual exemplar. One cyclist exemplar retrieves other cyclists.

3 · Existing evidence inventory

EvidenceAlready availableWhat it answersDecision
Architecture ablationFull 30.2 AP; without 2D head 11.1; without O2M 27.7; without geometry loss 28.5; without confidence 29.4; without deep supervision 29.9. Components contribute beyond simply attaching a 3D head.Paper evidence
Ready to draft
Real / sensor depthOmni3D with provided depth and ScanNet GT depth provide independent 3D-GT evidence. Stereo4D stereo depth (7.5→27.7 AP) and DROID FoundationStereo depth (6.8→24.1 AP) demonstrate real-depth modality transfer, but their 3D labels remain pipeline-derived. Optional depth works beyond the in-the-wild MoGe setting, while only Omni3D/ScanNet cleanly address label independence.Paper evidence
Ready to draft
Data mixture ablationOmni3D-only, Other-3D, human-only, synthetic-only, ITW-full, and full-mixture checkpoints were evaluated on Omni3D and ScanNet. The result is dataset-dependent and does not support a monotonic “more data is always better” claim; retain internally and avoid positive headline use.Completed locally
Pipeline / annotator analysisCandidate selection shares and rejection rates; VLM-human AUC 0.66, r=0.30, and 73.4% top-2 coverage. VLM scores correlate with human preferences; the new independent-GT replay now separately measures metric 3D localization. Paper evidence
Completed locally
Prompt modalitiesCorrected matched-protocol text, point, box, and exemplar evaluation is complete; point uses its point-trained checkpoint and click-aware selection. Six point/exemplar qualitative panels passed manual inspection. Directly tests whether all four claimed prompt modes work at inference. Completed locally

4 · Reviewer Z8Kd

ConcernPlanned responseEvidence sourceActionStatus
The promptable claim is broader than the evaluation: no point/exemplar quantitative or qualitative results. Report matched-protocol quantitative results for all four prompt modes and provide a point/exemplar qualitative grid. State each prompt's target contract explicitly. The paper establishes joint supervision; the corrected evaluation uses the corresponding prompt-trained checkpoint, one evaluator, and deterministic prompt construction. Corrected point/box/exemplar audit completed; full table and qualitative grid are above. Completed locally
WildDet3D-Bench may favor models trained with the same annotation pipeline. “Other 3D” not helping strengthens this concern. Acknowledge that this benchmark alone is not independent. Lead with the hidden-GT depth-source counterfactual, standard Omni3D results, and real-depth transfer. Do not claim monotonic benefits from every data mixture. The completed Omni3D/ScanNet mixture table is mixed and is retained internally rather than used as positive headline evidence. No rerun. Narrow the claim and use stronger independent evidence. Completed locally

5 · Reviewer qeMe

ConcernPlanned responseEvidence sourceActionStatus
On WildDet3D-Bench, MoGe-derived labels and MoGe depth input may create circularity. Agree that 22.6→41.6 alone is not independent, then report the new leakage-safe counterfactual: with fixed candidates and VLM selector, UniDepthV2 preserves coverage and improves mean oriented IoU from 6.92% to 12.58% and center-within-2m from 50.61% to 65.69%. Independent Omni3D hidden GT, plus existing Stereo4D and DROID real-depth transfer. Depth-source robustness experiment completed. Completed locally
Architecture, new data, and stronger pretrained backbones are confounded. Use the existing Omni3D-only comparison and component ablation. Use the independent data ablation to isolate the incremental effect of WildDet3D-Data. Avoid claiming that the entire gain comes from a novel backbone architecture. Architecture ablation + Omni3D-only model + new independent data table. No matched-backbone retraining in the four-day window. Paper evidence
Completed locally

6 · Reviewer bYkw

ConcernPlanned responseEvidence sourceActionStatus
Metric scale cannot be inferred uniquely from a single RGB image. Agree and revise “metric world knowledge” to “metric 3D estimates from learned visual, category, and scene priors.” Explicitly state that monocular metric scale is underdetermined and optional depth reduces this ambiguity. The method already states that monocular 3D is ill-posed, but later wording overreaches. Text revision only.Ready to draft
MoGe errors may propagate into labels; GPT size and axis-ratio filters may be unreliable under category variance and rotation. Report the completed hidden-GT Omni3D replay: raw candidate oracle, VLM selection on score-transferred raw geometry, deployed aligned geometry, and the UniDepthV2 counterfactual. Lead with center-distance recall and retain oriented IoU as a strict diagnostic. VLM-human correlation alone was insufficient; the new replay directly measures metric localization. The alternative estimator improves geometry without changing coverage. GT-backed pipeline-stage and depth-source audits complete.Completed locally
Technical novelty is limited and gains may come from foundation modules. Position the contribution as a unified prompt-conditioned 3D system with optional-depth residual correction plus a validated scaling pipeline. Cite existing component ablations and independent transfer; do not claim novelty for SAM3, DINOv2, or ControlNet-style fusion individually. Existing architecture ablation and independent benchmarks.Draft-only. Paper evidence
Ready to draft

7 · Area Chair / meta-review

ConcernPlanned responseEvidence sourceActionStatus
Dataset annotations may not be high quality. Provide direct metric validation against hidden independent 3D GT, not only annotator QC. Show which pipeline stages improve or hurt geometric quality and demonstrate that the conclusion does not depend on MoGe-2. Completed GT-backed pipeline validation and UniDepthV2 counterfactual, supported by existing human QC statistics. Report both strict IoU and loose distance metrics with interpretation boundaries. Completed locally
If technical novelty is limited, the work may look like dataset engineering rather than a NeurIPS contribution. Frame the scientific contribution around learning transferable open-world metric 3D representations from heterogeneous model-assisted supervision, tested on independent GT and across prompt/depth regimes. The architecture is intentionally simple and validated component-wise. Architecture ablation, prompt evaluation, independent data ablation, and transfer tables. Evidence synthesis; no large new backbone study.Paper evidence
Completed locally

8 · Experiment decision matrix

WorkstreamExact comparisonTraining?PriorityStatus
A · Independent data ablationOmni3D-only / +Other-3D / +ITW-human / +ITW-synthetic / +ITW-full / full mixture, evaluated on Omni3D and ScanNet with one config. No; reuse existing checkpoints.P0Completed locally
B0 · GT-box lifting auditMonocular vs GT-depth lifting on all six Omni3D test sets, direct oriented IoU and geometric errors.No; saved predictions. P0Completed locally
B · GT-backed pipeline validationRaw five-generator candidates → alignment/filtering → VLM selection, evaluated against hidden 3D GT. No new Prolific choices were collected on this replay subset; human-filter evidence comes from the paper's 481K-candidate audit. No model training; pipeline replay.P0Completed locally
C · Prompt evaluationText / box / point / exemplar. Report official Omni3D 3D AP metrics and point/exemplar qualitative examples.No; reuse final checkpoint. P1Completed locally
D · Depth-source robustnessMoGe-2 → UniDepthV2 on the same 300-image hidden-GT replay, holding camera, candidate outputs, and VLM selector fixed. No training; one 1-GPU generation stage plus CPU/GPU replay stages. P0Completed locally
E · Independent annotation auditTwo blinded expert raters score 200–300 random annotations for center/depth, dimensions, orientation, and usability; report 95% CI and inter-rater agreement.No GPU; requires human raters. Optional P1Ready to draft
Matched-backbone 3D-MOODNot scheduled.Would require expensive retraining and code adaptation.SkipReady to draft
Additional depth estimatorsNot needed after the conclusive MoGe-2 versus UniDepthV2 result.NoSkipCompleted locally

9 · Final four-day decision

DecisionReason
No more GPU experiments recommended.The four-prompt evaluation, hidden-GT replay, UniDepthV2 counterfactual, architecture ablations, and real-depth benchmarks already answer the actionable reviewer questions. Another depth estimator has low marginal value.
Do not launch matched-backbone retraining.The AC says architectural simplicity is not itself disqualifying; multi-day retraining is high-risk and less relevant than the now-completed independence test.
Run one human audit only if raters are immediately available.A blinded, independent 200–300-box audit directly strengthens the human-verification claim and needs no GPU. If raters are unavailable, submit with the completed hidden-GT audit instead of inventing a weak proxy.
Spend the remaining window on writing.Lead with the three pillars, answer circularity using the UniDepth result, explain the point/box information gap, retain strict metric limitations, and remove claims unsupported by the mixed data-mixture rerun.
Maintained rebuttal evidence page. New experiment rows include exact Beaker provenance; paper-evidence rows are transcribed from the submission snapshot. The remaining work is rebuttal writing and, only if independent raters are available, one blinded human audit.