ViewMind3D: 모듈형 시야 인지 추론을 활용한 학습 불필요 3차원 질의응답
ViewMind3D: Modular View-Aware Inference for Training-Free 3D-QA
최근 대규모 언어 모델(LLM)과 시각-언어 모델(VLM)의 발전은 자율 에이전트 및 로봇 인식에 중요한 기능을 제공하는 3차원 질의응답(3D-QA) 분야에서 새로운 가능성을 열었습니다. 그러나 대부분의 기존 방법은 비용이 많이 드는 어노테이션을 사용한 3차원 특화 학습 또는 미세 조정을 필요로 하며, 이는 확장성과 실제 적용 가능성을 제한합니다. 본 논문에서는 완전한 3차원 재구성이 필요 없는 장면의 다중 시점 관찰에 대한 모듈형 프레임워크인 extbf{ViewMind3D}를 제시합니다. 이 프레임워크는 3D-QA 작업을 다음 네 가지 해석 가능한 구성 요소로 분해합니다: (1) 질문 기반 다중 시점 선택, (2) 언어 조건부 객체 단서를 사용한 가이드형 시각적 정렬, (3) 버드아이뷰(BEV) 관점을 통한 공간 맥락 인코딩, 그리고 (4) 역할 기반 추론을 통한 구조화된 답변 생성. 이러한 설계는 모델 튜닝 없이도 구조적이고 견고하며 해석 가능한 추론을 가능하게 합니다. ScanQA 및 SQA3D 데이터셋에 대한 실험 결과에서 ViewMind3D는 기존의 학습 불필요한 3D-LLM과 미세 조정된 3D-LLM과 경쟁력 있는 성능을 달성했습니다. 특히, 본 방법은 SQA3D의 공간적 정렬 질문 유형(예: "무엇" 질문)에서 성능이 향상되었으며, 전반적인 정확도는 50.8%로 높고 ScanQA에서 CIDEr 점수는 73.4점을 기록했습니다. 이러한 결과는 일반적인 LLM 및 VLM을 모듈 방식으로 활용하여 실제 환경에서의 로봇 인식을 위한 효과적인 3차원 추론이 가능하다는 것을 보여줍니다.
Recent advances in large language models (LLMs) and vision-language models (VLMs) have enabled new possibilities for 3D question answering (3D-QA), a key capability for embodied AI and robotic perception. However, most existing methods rely on 3D-specific training or fine-tuning with costly annotations, limiting their scalability and real-world applicability. We present \textbf{ViewMind3D}, a fully training-free and modular framework for 3D spatial reasoning over multi-view observations of a scene without requiring complete 3D reconstruction. The framework decomposes the 3D-QA task into four interpretable components: (1) question-driven multi-view selection, (2) guided visual grounding with language-conditioned object cues, (3) spatial context encoding via a bird's-eye-view (BEV) viewpoint indicator, and (4) structured answer generation through role-based reasoning. This design enables structured, robust, and interpretable reasoning without requiring model tuning. Experimental results on ScanQA and SQA3D show that ViewMind3D achieves competitive performance compared to prior training-free and fine-tuned 3D-LLMs. In particular, our method improves performance on spatially grounded question types, such as ``What'' questions in SQA3D, while maintaining strong overall accuracy (50.8\%) and achieving 73.4 CIDEr on ScanQA. These results demonstrate that effective 3D reasoning can be achieved through modular orchestration of general-purpose LLMs and VLMs for robotic perception in real-world environments.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.