2603.08131v2 Mar 09, 2026 cs.RO

UniGround: 학습 없이 장면 분석을 통해 구현하는 범용 3차원 시각적 객체 지칭

UniGround: Universal 3D Visual Grounding via Training-Free Scene Parsing

Yu Fang
Yu Fang
Citations: 8
h-index: 2
Weisheng Xu
Weisheng Xu
Citations: 5
h-index: 2
Renjing Xu
Renjing Xu
Citations: 20
h-index: 2
Jiaxin Zhang
Jiaxin Zhang
Citations: 11
h-index: 2
Yunheng Wang
Yunheng Wang
Citations: 12
h-index: 1
Wei Lu
Wei Lu
Citations: 22
h-index: 3
Taowen Wang
Taowen Wang
Citations: 131
h-index: 4
Shuning Zhang
Shuning Zhang
Citations: 14
h-index: 2
Yixiao Feng
Yixiao Feng
Citations: 12
h-index: 1

3차원 시각적 객체 지칭(3DVG)은 3차원 환경에서 자연어 설명을 기반으로 객체를 위치시키는 기술로, 로봇 인공지능 응용 분야의 핵심입니다. 기존 모델들은 개방형 어휘 추론을 가능하게 하지만, 일반적으로 사전 생성된 후보 객체에 의존하여 두 가지 병목 현상을 야기합니다. 첫째, 데이터셋 특화된 3차원 제안 모델이 분포 변화로 인해 목표 객체를 놓치거나, 일부만 감지하거나, 잘못 그룹화하여 VLM(Visual Language Model) 추론에서 제외되는 '후보 객체 병목' 현상이 발생합니다. 둘째, 불완전한 시각 정보로 인한 '증거 병목' 현상도 존재합니다. 전체 렌더링은 공간적 맥락을 유지하지만 객체의 세부 사항을 가리고, 후보 객체를 중심으로 한 시점은 로컬 표현은 잘 잡아내지만 전역적인 맥락이 부족합니다. 이러한 병목 현상을 해결하기 위해, 우리는 학습 없이 동작하는 3DVG 프레임워크인 UniGround를 제안합니다. UniGround는 글로벌 후보 필터링(Global Candidate Filtering)과 문맥 기반 정밀 지칭(Contextual Precision Grounding)을 통해 두 가지 병목 현상을 모두 해결합니다. 글로벌 후보 필터링은 데이터셋 학습이 필요 없는 3차원 감지 모델, 특정 작업에 대한 제안 지도 또는 미리 정의된 박스 및 범주 정보를 사용하지 않고, 3차원 구조와 다중 시점 의미 단서를 통해 위상적으로 일관적인, 클래스 불변의 후보 객체를 생성합니다. 문맥 기반 정밀 지칭은 글로벌 공간적 맥락과 후보 객체 중심의 시각 정보에 대해 동시에 추론하고, 신뢰성 있는 목표 객체 식별을 위해 폐루프 일관성 검증을 수행합니다. UniGround는 ScanRefer 데이터셋에서 46.1%/34.1%의 Acc@0.25/0.5를 달성했으며, EmbodiedScan의 ARKitScenes 서브셋에서는 28.7%의 Acc@0.25를 달성했습니다. 추가적인 실험을 통해 데이터셋 특화된 3차원 정보 없이도 경쟁력 있는 지칭 성능을 보이며, 학습하지 않은 실내 환경으로의 일반화 능력과 실제 재구성 과정에서 발생하는 노이즈 및 다양한 도메인 변화에 대한 강건성을 입증합니다.

Original Abstract

3D Visual Grounding (3DVG) localizes objects from natural-language descriptions in 3D scenes and is fundamental to embodied AI applications. Although foundation models enable open-vocabulary reasoning, they typically rely on pre-generated candidates, creating two sequential bottlenecks. The \emph{candidate bottleneck} occurs when dataset-specific 3D proposal models miss, fragment, or incorrectly group targets under distribution shifts, excluding them from VLM reasoning. The \emph{evidence bottleneck} stems from incomplete visual evidence: global renderings preserve spatial context but obscure object details, whereas candidate-centric views capture local appearance but lack global context. To address these bottlenecks, we propose UniGround, a zero-shot 3DVG framework that addresses both bottlenecks through Global Candidate Filtering and Contextual Precision Grounding. Global Candidate Filtering constructs topology-consistent, class-agnostic candidates from 3D topology and multi-view semantic cues, without dataset-trained 3D detectors, task-specific proposal supervision, or predefined box and category priors. Contextual Precision Grounding jointly reasons over global spatial context and candidate-centric visual evidence, followed by closed-loop consistency verification for reliable target identification. UniGround achieves 46.1\%/34.1\% Acc@0.25/0.5 on ScanRefer and 28.7\% Acc@0.25 on the evaluated ARKitScenes subset of EmbodiedScan. Further experiments demonstrate competitive grounding without dataset-specific 3D priors, cross-dataset generalization to unseen indoor scenes, and robustness to real-world reconstruction noise and practical domain shifts.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!