2607.24407v1 Jul 27, 2026 cs.CV

사고 토큰 혼합(Mixture-of-Thought-Tokens): 자유 형식의 다중 모드 정렬을 위한 인식 및 추론 통합

Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding

Kongming Liang
Kongming Liang
Citations: 12
h-index: 1
Xin Wei
Xin Wei
Citations: 23
h-index: 1
Hao Li
Hao Li
Citations: 313
h-index: 5
Hongbo Sun
Hongbo Sun
Citations: 31
h-index: 3
Jingmin Xin
Jingmin Xin
Citations: 0
h-index: 0
Han Fang
Han Fang
Citations: 95
h-index: 5
Jinglin Xu
Jinglin Xu
Citations: 41
h-index: 3
Ye Yuan
Ye Yuan
Citations: 26
h-index: 2
Hao Sun
Hao Sun
Citations: 24
h-index: 1
Tianyi Gao
Tianyi Gao
Citations: 16
h-index: 2
Tianyi Ding
Tianyi Ding
Citations: 0
h-index: 0
Xiaodong Dong
Xiaodong Dong
Citations: 4
h-index: 1

다중 모드 대규모 언어 모델은 정렬 작업에서 상당한 발전을 이루었지만, 기존 방법은 여전히 정확한 위치 파악과 복잡한 추론을 통합하는 데 어려움을 겪고 있습니다. 텍스트 기반 방법은 좌표 또는 인덱스 예측에 의존하여 모델의 밀집된 시각적 객체에 대한 인식 능력을 심각하게 제한합니다. 반면, 잠재 토큰 기반 방법은 고유한 공간 참조가 없는 특수 토큰을 사용하며, 사고 단계를 포함하지 않는 디코딩 메커니즘을 사용하여 고급 추론 능력을 약화시킵니다. 결과적으로, 인식과 추론 모두에서 뛰어난 성능을 보이는 통합 프레임워크를 개발하는 것은 여전히 어려운 과제입니다. 이러한 문제를 해결하기 위해, 우리는 인식-추론 간의 격차를 해소하고 MLLM이 다양한 유형의 정렬 쿼리에 활용될 수 있도록 하는 새로운 자유 형식의 다중 모드 정렬 방법인 Mixture-of-Thought-Tokens (Motto)를 제안합니다. 구체적으로, 우리는 명확한 공간적 대응성과 시각적 해석 가능성을 위해 특수 토큰을 공간 위치와 명시적으로 연결하는 공간 기반 사고 토큰화(Spatially-Grounded Thought Tokenization)를 도입했습니다. 또한, 우리는 다양한 복잡성을 가진 작업에서 강력한 정렬 성능을 달성하기 위해, 상호 연결된 추론 체인 내에서 정렬 모드를 동적으로 전환하는 컨텍스트 적응형 체인-오브-토큰(Context-Adaptive Chain-of-Tokens)을 설계했습니다. 또한, 인식-추론 간의 격차를 평가하기 위한 새로운 참조 표현 이해 벤치마크인 PR-Bench를 구축했습니다. 광범위한 실험 결과, Motto는 다양한 자유 형식 정렬 작업에서 최고 수준의 성능을 달성하는 것으로 나타났습니다.

Original Abstract

Multimodal Large Language Models have made great progress in grounding tasks, yet existing methods still struggle to unify precise localization and complex reasoning. For one thing, text-based methods rely on coordinates or index prediction, severely limiting the perceptual capabilities of the model for dense visual objects. Meanwhile, latent token-based methods employ special tokens without inherent spatial references and use a decoding mechanism that lacks thinking steps, weakening high-level reasoning capabilities. Consequently, developing a unified framework that excels in both perception and reasoning remains challenging. To address this, we propose Mixture-of-Thought-Tokens (Motto), a new free-form multimodal grounding method that bridges the perception-reasoning gap, enabling MLLMs to empower diverse, arbitrary grounding queries. Specifically, we introduce Spatially-Grounded Thought Tokenization to explicitly align special tokens with spatial locations for clear spatial correspondence and visual interpretability. We further design a Context-Adaptive Chain-of-Tokens that dynamically switch grounding modes within an interleaved reasoning chain, achieving robust grounding across tasks of varying complexity. In addition, we construct PR-Bench, a new referring expression comprehension benchmark to evaluate the perception-reasoning gap. Extensive experiments demonstrate that Motto achieves state-of-the-art performance across diverse free-form grounding tasks.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!