SPARGen: 멀티모달 생성 방식을 통한 공간 인식 및 추론의 통합
SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation
시각적 관찰로부터의 공간 인식 및 추론은 기하학적 구조 복원, 대응 관계 설정, 그리고 공간 관계 이해를 필요로 합니다. 기존 접근 방식들은 일반적으로 이러한 능력을 별도로 처리하며, 작업별 아키텍처나 외부 기하학 모듈을 사용합니다. 이는 동일한 물리적 장면의 상호 보완적인 표현 간의 지식 전달을 제한합니다. 본 논문에서는 3D 재구성, 밀집 대응 관계 찾기, 그리고 공간 추론을 명령 기반 생성 작업으로 통합하는 멀티모달 프레임워크인 SPARGen을 소개합니다. SPARGen은 구조화된 데이터와 언어적 출력을 압축하여 토큰 시퀀스로 표현하고, 동시에 이미지에 정렬된 형태의 밀집적인 기하학적 필드를 생성합니다. 이를 통해 공간적 감독 신호가 원활하게 공유되는 표현을 형성하며, 단일 멀티모달 생성 모델 내에서 학습이 가능합니다. 3D 재구성, 대응 관계 찾기 및 공간 추론에 대한 다양한 벤치마크 실험 결과는 SPARGen이 단일의 통합된 멀티모달 생성 프레임워크 내에서 다양한 공간 관련 작업에서 경쟁력 있는 성능을 달성함을 보여줍니다.
Spatial perception and reasoning from visual observations require recovering geometric structure, establishing correspondences, and understanding spatial relations. Existing approaches typically address these capabilities separately using task-specific architectures or external geometric modules, limiting knowledge transfer among complementary representations of the same physical scene. We introduce SPARGen, a unified multimodal framework that casts 3D reconstruction, dense correspondence, and spatial reasoning as instruction-conditioned generation tasks. SPARGen serializes compact structured and linguistic outputs as token sequences while generating dense geometric fields in image-aligned forms, enabling spatial supervision to jointly shape shared representations within a native multimodal generative model. Experiments across benchmarks for 3D reconstruction, correspondence, and spatial reasoning show that SPARGen achieves competitive performance across heterogeneous spatial tasks within a single native multimodal generative framework.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.