프롬프트가 픽셀로 변할 때: 다중 모드 추론을 위한 프롬프트-영역 연계
When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning
다중 모드 대규모 언어 모델은 점점 더 스크린샷과 문서에 대한 추론을 수행하며, 이때 작업 자체가 픽셀 단위로 표현될 수 있습니다. 그러나 현재의 평가 지표는 대부분 질문을 텍스트 형태로 제시하므로, 모델이 다양한 입력 채널에서 동일한 명령어를 얼마나 효과적으로 사용하는지 불분명합니다. 본 연구에서는 시각적 작업 의미(Visualized Task Semantics, VTS)라는 제어된 방법을 도입하여 질문을 이미지 내부로 이동시키면서 문제와 정답은 그대로 유지했습니다. 6개의 MLLM 모델과 4가지 평가 지표를 사용하여 실험한 결과, 모든 24개의 모델-작업 쌍에서 정확도가 평균적으로 17.8 포인트 감소했습니다. 모델들은 시각적인 질문을 대체로 정확하게 인식하지만, 이를 활용하는 데 실패하는 경우가 많았으며, 이는 광학 문자 인식(OCR) 이상의 의미론적 채널 격차를 드러냅니다. 이러한 격차를 줄이기 위해, 프롬프트-영역 연계 방법을 제안합니다. 이 방법은 질문 영역을 텍스트 의미와 일치시키고, 가려진 시점에서 해당 영역의 명확한 표현을 복원하는 것이 핵심입니다. 제안된 방법은 기존 방식과 동일한 학습 비용으로 VTS 정확도를 58.0에서 66.3으로 향상시켰으며, 원래 인터페이스에서의 정확도를 유지하고 추론 과정에서 OCR이나 영역 메타데이터가 필요하지 않습니다. 작업 관련 텍스트를 읽고 이를 추론을 위한 명령어로 활용하는 것은 별개의 능력입니다.
Multimodal large language models increasingly reason over screenshots and documents where the task itself may be written in pixels. Yet benchmarks usually place questions in text, leaving it unclear whether models use the same instruction equally well across channels. We introduce Visualized Task Semantics (VTS), a controlled intervention that moves the question into the image while keeping the source problem and answer fixed. Across six MLLMs and four benchmarks, accuracy drops in all 24 model-task pairs, by 17.8 points on average. Models often transcribe the visual question correctly yet fail to use it, exposing a semantic channel gap beyond OCR. To reduce this gap, we present prompt-region grounding, whose core design aligns the question region with typed semantics and recovers its clean representation from a masked view. At matched training cost, our method raises four-benchmark VTS accuracy from 58.0 to 66.3 while preserving accuracy on the original interface, and requires no OCR or region metadata at inference. Reading task-bearing text and grounding it as an instruction for reasoning are distinct capabilities.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.