SSR3D-LLM: 잠재 단계 기반 구조화된 공간 추론을 통한 통합 3D-LLM에서의 정밀 객체 지시
SSR3D-LLM: Structured Spatial Reasoning via Latent Steps for Fine-Grained Grounding in Unified 3D-LLMs
3차원 객체 지시는 자연어 설명을 바탕으로 3차원 장면 내에서 특정 객체를 찾아내는 기술입니다. 통합 인스턴스 중심 3D-LLM은 지시 작업 외에도 대화, 질의 응답, 이미지 설명 등 다양한 작업을 수행하는 것을 목표로 하지만, 많은 모델들이 관계적 명령을 하나의 선택으로 압축하는 단일 포인터 방식의 지시 결정을 사용합니다. 이는 문맥 객체와 공간 관계를 통해 여러 동일 클래스 후보를 배제해야 하는 정밀한 쿼리에서 취약점을 드러냅니다. 본 논문에서는 통합 3D-LLM을 위한 구조화된 지시 인터페이스인 Structured Spatial Reasoning 3D-LLM (SSR3D-LLM)을 제안합니다. SSR3D-LLM은 주어진 Mask3D 객체 후보에 대해, LLM이 입력 쿼리로부터 생성된 잠재적인 공간 추론 단계와 메모리 토큰 시퀀스를 작성하고, 기하학적 정보를 고려하는 평가기가 이러한 잠재 단계를 순차적으로 읽어 들여 스테핑-길이 마스킹을 통해 후보 순위를 점진적으로 개선합니다. 잠재 단계는 표준 벤치마크 목표 감독 학습과 함께 보조적인 참조 힌트 감독 학습을 통해 학습되며, 추론 과정에서는 입력 쿼리와 Mask3D 제안만 사용됩니다. 실험 결과, SSR3D-LLM은 ReferIt3D, ScanRefer, Multi3DRef 데이터셋에서 통합 3D-LLM 기반 모델 중 가장 뛰어난 성능을 보였으며, 정밀한 지시 작업에서는 단일 포인터 QPG 모델보다 상당한 개선 효과를 나타냈습니다. 또한 기존의 통합 3D-LLM에 비해 일관된 성능 향상을 보여주면서도 기본적인 언어 처리 기능을 유지합니다.
3D object grounding localizes referred objects in a 3D scene from natural language. Unified instance-centric 3D-LLMs aim to solve grounding together with dialog, QA, and captioning, yet many rely on a single pointer-style grounding decision that compresses a relational instruction into one selection. This is brittle for fine-grained queries where multiple same-class candidates must be ruled out by context objects and spatial relations. We propose Structured Spatial Reasoning 3D-LLM (SSR3D-LLM), a structured grounding interface for unified 3D-LLMs. Given fixed Mask3D object proposals, the LLM writes a sequence of latent spatial reasoning steps and memory tokens from the query, and a geometry-aware scorer reads these latent steps in order to refine candidate rankings step by step with step-length masking. The latent steps are learned from standard benchmark target supervision with auxiliary referential-cue supervision during training, while inference uses only the input query and Mask3D proposals. Across ReferIt3D, ScanRefer, and Multi3DRef, SSR3D-LLM achieves the strongest results among unified 3D-LLM baselines, with substantial gains over the single-pointer QPG baseline on fine-grained grounding and consistent improvements over prior unified 3D-LLMs, while preserving the default language-task route.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.