OrientSAM: 방향 인지 기반 공간 정렬을 통한 다중 모드 시각적 추론에서 발생하는 카메라 중심 편향 완화
OrientSAM: Mitigating Camera-Centric Shortcut in Multimodal Spatial Reasoning via Orientation-Aware Spatial Alignment
다중 모드 대규모 언어 모델(MLLM)은 여전히 원근 변환이 필요한 공간 추론에 어려움을 겪습니다. 특히, 이러한 모델들은 종종 참조 객체의 관점보다는 카메라 중심적인 단서에 의존하여 작동하며, 이는 비-카메라 참조 환경에서 체계적인 오류를 야기합니다. 본 논문에서는 이러한 실패 사례를 분석하고, 객체 방향이 카메라 중심 편향 행동의 핵심 요인임을 보여줍니다. 이 문제를 해결하기 위해, 우리는 방향 인지 기반 공간 정렬 프레임워크인 OrientSAM을 제안합니다. OrientSAM은 방향 인지 토큰과 푸리에 기반 각도 인코딩을 통해 다중 모드 표현에 명시적인 방향 정보를 주입하고, 또한 커리큘럼 학습 전략을 채택하여 원근 변환에 대한 추론 능력을 점진적으로 향상시킵니다. 더불어, 대규모 이미지로부터 방향 인지 공간적 감독 신호를 생성하기 위한 데이터 구축 파이프라인을 개발했습니다. Spatial-MM, ViewSpatial 및 3DSRBench에서의 실험 결과는 OrientSAM이 강력한 기준 모델보다 일관되게 우수한 성능을 보임을 보여주며, 특히 카메라 시점이 아닌 환경, 사람 중심적인 작업, 그리고 방향에 민감한 작업에서 두드러진 성능 향상을 나타냅니다. 이러한 결과는 명시적인 방향 모델링이 카메라 중심 편향을 완화하고 다중 모드 모델의 보다 강력한 위치 기반 공간 추론을 가능하게 하는 데 중요하다는 것을 입증합니다.
Multimodal large language models (MLLMs) still struggle with spatial reasoning that requires perspective transformation. In particular, they often rely on camera-centric cues rather than reasoning from the reference object's viewpoint, leading to systematic errors in non-camera reference settings. In this paper, we first analyze this failure mode and show that object orientation is a key factor underlying such camera-centric shortcut behavior. To address this issue, we propose OrientSAM, an orientation-aware spatial alignment framework for multimodal models. OrientSAM injects explicit orientation information into multimodal representations through orientation-aware tokens and Fourier-based angle encoding, and further adopts a curriculum learning strategy to progressively improve perspective-aware reasoning. In addition, we build a spatial data construction pipeline to generate orientation-aware spatial supervision from large-scale images. Experiments on Spatial-MM, ViewSpatial, and 3DSRBench show that OrientSAM consistently outperforms strong baselines, especially on non-camera-view, person-centric, and orientation-sensitive tasks. The results further demonstrate that explicit orientation modeling is important for mitigating camera-centric shortcut behavior and enabling more robust allocentric spatial reasoning in multimodal models.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.