2608.02980v1 Aug 04, 2026 cs.CV

Qwen-3D: 공간 인지 능력을 위한 범용적인 3차원 시각-언어 모델

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Katerina Fragkiadaki
Katerina Fragkiadaki
Citations: 4,306
h-index: 32
Ayush Jain
Ayush Jain
Citations: 39
h-index: 3
Yifan Liu
Yifan Liu
Citations: 4
h-index: 1

대규모 다중 모드 모델(LMM)은 이미지 및 짧은 동영상에서 뛰어난 성과를 거두었지만, 프레임 중심 토큰화와 제한된 컨텍스트 창으로 인해 긴 동영상을 처리하는 데는 어려움이 있습니다. 3차원 기하학 정보는 시각 데이터 스트림을 압축하는 자연스러운 메커니즘을 제공합니다. 깊이 정보와 카메라 자세를 통해 다양한 시점과 시간 단계를 하나의 일관된, 세계 좌표계에 정렬된 표현으로 결합할 수 있습니다. 최근의 3차원 LMM은 공간 추론 능력을 향상시키기 위해 기하학 정보를 활용하지만, 여전히 객체 인식 및 분할 작업에서는 전문적인 3차원 인지 시스템보다 성능이 뒤쳐집니다. 우리는 이러한 한계의 주요 원인이 기하학 정보에 기반한 디코딩 과정이라는 점을 지적합니다. 기존 방법은 3차원 예측 결과를 언어 토큰, 제안 선택 또는 간단한 객체 지칭 질의를 통해 전달함으로써 언어 추론과 정밀한 기하학적 예측 간의 병목 현상을 초래합니다. 이러한 문제점을 해결하기 위해, 우리는 Qwen-3D를 소개합니다. Qwen-3D는 다중 시점 기하학적 정보를 사용하여 Qwen 모델 내에서 시각 정보를 압축하는 기하학 정보 기반 LMM으로, 정적인 장면에서의 효율적인 장기 시각 추론을 가능하게 합니다. Qwen-3D는 시각 토큰에 3차원 로터리 위치 임베딩을 추가하여 어텐션 메커니즘이 독립적인 이미지 프레임 간이 아니라 직접 3차원 공간에서 작동하도록 함으로써 확장 가능한 교차 시점 및 시간 추론을 용이하게 합니다. 언어 정보와 기하학 정보를 연결하기 위해, Qwen-3D는 언어를 기본 3차원 장면 표현에 직접 연결하는 질의 기반 분할 디코더를 통합하여 이미지 및 동영상에서 객체 지칭, 인스턴스 분할 및 시각적 질문 응답을 통일합니다. 다양한 벤치마크 테스트 결과, Qwen-3D는 기존의 3차원 LMM보다 뛰어난 성능을 보이며, 여러 대규모 독점적인 2차원 모델보다 우수한 결과를 나타냅니다. 특히, Qwen-3D는 2차원 및 3차원 데이터를 함께 학습함으로써 표준 2차원 시각-언어 벤치마크에서도 강력한 성능을 유지하며 이러한 개선 효과를 달성합니다.

Original Abstract

Large Multimodal Models (LMMs) have achieved remarkable success on images and short videos, yet scaling them to long videos remains challenging due to frame-centric tokenization and limited context windows. 3D geometry provides a natural compression mechanism for visual streams: depth and camera pose enable observations from multiple views and time steps to be fused into a persistent, world-aligned representation. While recent 3D LMMs leverage geometry-aware representations to improve spatial reasoning, they continue to lag behind specialist 3D perception systems on grounding and segmentation tasks. We argue that a key limitation is geometry-aware decoding: existing methods communicate 3D predictions through language tokens, proposal selection, or lightweight grounding queries, creating a bottleneck between language reasoning and dense geometric prediction. Building on these insights, we introduce Qwen-3D, a geometry-aware LMM that compresses visual information within the Qwen backbone using multi-view geometric cues, enabling efficient long-horizon visual reasoning over static scenes. Qwen-3D augments visual tokens with 3D Rotary Positional Embeddings, allowing attention to operate directly in 3D scene space rather than across independent image frames and thereby facilitating scalable cross-view and temporal reasoning. To bridge language and geometry, Qwen-3D incorporates a query-based segmentation decoder that grounds language directly in the underlying 3D scene representation, unifying referential grounding, instance segmentation, and visual question answering across both images and videos. Across a diverse set of benchmarks, Qwen-3D surpasses existing 3D LMMs and outperforms several large proprietary 2D models. Notably, Qwen-3D achieves these improvements while maintaining strong performance on standard 2D vision-language benchmarks by jointly training on 2D and 3D data.

0 Citations
0 Influential
16 Altmetric
80.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!