Linguistically Diverse Diagnostic Benchmark for 3D Visual Grounding
This benchmark reframes evaluation of 3D visual grounding around the linguistic diversity of referring prompts rather than raw dataset scale. The authors define a framework that categorizes grounding prompts by linguistic phenomena (such as negation, coarse-grained references, anaphoric reference resolution, and varied syntax) and assemble a compact diagnostic dataset that deliberately spans these patterns. Students learn how a targeted, linguistically-analyzed evaluation set exposes failure modes of open-vocabulary grounding models on out-of-distribution language that large homogeneous benchmarks miss, and how to reason about coverage of natural-language variation when designing vision-language benchmarks.
ViGiL3D: A Linguistically Diverse Dataset for 3D Visual Grounding Austin T. Wang1 ZeMing Gong1
This computer-vision and computational-linguistics paper addresses 3D visual grounding, the task of localizing the entities in a 3D scene that are referred to by a natural-language description. Obser…