2606.19965v1 Jun 18, 2026 cs.CV

ROSE: 다중 모드 모델의 인지-행동 격차 벤치마킹

ROSE: Benchmarking the Perception-to-Action Gap in Multimodal Models

Yihao Wang
Yihao Wang
Citations: 50
h-index: 4
Keze Wang
Keze Wang
Citations: 873
h-index: 16
Zijian He
Zijian He
Citations: 155
h-index: 4
Jie Ren
Jie Ren
Citations: 368
h-index: 5

다중 모드 대규모 언어 모델(MLLM)은 점차 시각 정보를 활용하여 행동하도록 요구받고 있지만, 동일한 장면이라도 서로 다른 작업 맥락에서 다른 행동이 필요할 수 있습니다. 모델이 동일한 시각적 증거를 사용하여 현재 맥락에 맞는 행동을 얼마나 안정적으로 수행할 수 있을까요? 이 질문에 답하기 위해, 우리는 시각 장면은 고정하고 영역 제약 조건과 필요한 기호 출력을 다양하게 변화시키는 통제된 벤치마크인 extsc{ROSE} ( extbf{R}eference-conditioned extbf{O}ddity and extbf{S}ymbolic extbf{E}xecution)를 소개합니다. extsc{ROSE}는 결합된 계수 및 좌표-행동 작업을 통해 모델이 암묵적인 다수 기준을 추론하고 변화하는 맥락에서 결과적으로 생성되는 세분화된 시각적 증거에 따라 행동할 수 있는지 테스트합니다. 최근 9개의 MLLM을 대상으로 실험한 결과, 모델의 성능은 계수에 중점을 둔 작업에서 영역 제약 조건이 있는 행동으로 전환될 때 최대 44.5%까지 감소했습니다. 인간 수행 능력은 98.8%에 달했지만, 동일한 모델이 올바른 개수를 반환하는 쌍을 이루는 장면과 영역에서도 이러한 격차가 지속됩니다. 또한 전역 클릭 및 일치된 지역 제어를 통해 좌표 매핑이 성능 저하의 일부를 설명할 뿐임을 보여주며, 공유된 시각적 증거를 맥락에 특화된 행동으로 변환하는 데 있어 모델 의존적인 고유한 병목 현상이 존재한다는 것을 밝혀냅니다.

Original Abstract

Multimodal large language models (MLLMs) are increasingly expected to act on visual information, yet the same scene may require different actions under different task contexts. How reliably can a model turn the same visual evidence into the action required by the current context? To answer this question, we introduce \textsc{ROSE} (\textbf{R}eference-conditioned \textbf{O}ddity and \textbf{S}ymbolic \textbf{E}xecution), a controlled benchmark that holds the visual scene fixed while varying region constraints and required symbolic outputs. Through coupled counting and coordinate-action tasks, \textsc{ROSE} tests whether models can infer an implicit majority reference and act on the resulting fine-grained visual evidence under changing contexts. Across nine recent MLLMs, performance drops by as much as 44.5 percentage points from counting-oriented tasks to region-conditioned action, despite 98.8\% human performance. The gap persists on paired scenes and regions for which the same model returns the correct count, while global-click and matched local controls show that coordinate grounding explains only part of the loss, revealing a distinct, model-dependent bottleneck in turning shared visual evidence into context-specific actions.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!