SpatioLM: 비전-언어 모델에서 일반적인 물리적 공간 지능을 향한 연구
SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models
비전-언어 모델(VLMs)은 상식 추론 작업에서는 뛰어난 성능을 보이지만, 시각적 공간 추론 능력은 부족합니다. 기존의 대부분 해결책들은 추가적인 3차원 사전 정보 입력 또는 외부 공간 인코더를 사용하는데, 이는 복잡성을 증가시키고 공간 미세 조정 이후 기본 VLM의 범용 능력을 저하시킵니다. 이러한 문제를 해결하기 위해, 저희는 추가적인 3차원 사전 정보 입력이나 외부 공간 인코더 없이 공간 지능을 향상시키는 파라미터 효율적인 extit{ extbf{Spatio}-vision extbf{L}anguage extbf{M}odels (SpatioLM)을 제안합니다. 구체적으로, 저희는 VLM에 내재된 공간적 지식을 활용하는 플러그 앤 플레이 방식의 비침습적 공간-비전 모듈을 설계했습니다. 또한, 모델이 물리적으로 일관성 있는 표현을 학습하도록 유도하기 위해, 가상 깊이 정보와 카메라 정보를 지도(supervision)로 활용합니다. 광범위한 실험 결과는 SpatioLM이 다양한 작업에서 상당한 성능 향상을 달성했으며, 특히 공간 인식 및 이해 능력 향상과 함께 일반적인 능력이 저하되는 것을 효과적으로 제한한다는 것을 보여줍니다. 주목할 만하게도, 저희 모델은 VSI-Bench에서 71.6이라는 인상적인 점수를 기록했습니다(70점을 넘는 최초의 모델). 또한, 로봇 제어와 같은 임베디드 조작 작업으로 전이했을 때에도 경쟁력 있는 성능을 보였습니다. 관련 코드는 다음 링크에서 확인하실 수 있습니다: [https://github.com/xiaomi-research/spatio-lm](https://github.com/xiaomi-research/spatio-lm)
Vision-Language Models (VLMs) perform well on commonsense reasoning tasks but struggle with visual spatial reasoning. Most existing solutions introduce extra 3D prior inputs or external spatial encoders, which increase complexity and degrade the underlying VLMs' general-purpose capabilities after spatial fine-tuning. To this end, we propose a parameter-efficient \textit{\textbf{Spatio}-vision \textbf{L}anguage \textbf{M}odels (SpatioLM)}, that enhances spatial intelligence without extra 3D prior inputs or third-party spatial encoders. Concretely, we design a plug-and-play and non-invasive spatio-vision module that elicits the spatial knowledge inherent in VLMs. Furthermore, we innovatively leverage pseudo depth and camera information as supervision to guide the model in learning physically coherent representations. Extensive experiments show that SpatioLM achieves significant improvements in diverse tasks, including spatial perception and understanding while effectively limiting the degradation of general capabilities. Notably, the model achieves an impressive score of 71.6 on the VSI-Bench (the first model to surpass 70). In addition, it attains competitive performance when transferred to embodied manipulation tasks. Code is available at \href{https://github.com/xiaomi-research/spatio-lm}{\faGithub~spatio-lm}.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.