2606.09585v1 Jun 08, 2026 cs.AI

광학적 추론: 텍스트를 넘어 이미지 자체를 표현력 있는 추론 매개체로 재해석

Optical Reasoning: Rethinking Images as an Expressive Reasoning Medium Beyond Text

Dongjie Cheng
Dongjie Cheng
Citations: 4
h-index: 1
Wenjie Li
Wenjie Li
Citations: 23
h-index: 3
Yongqing Li
Yongqing Li
Citations: 260
h-index: 7
Heming Xia
Heming Xia
Citations: 421
h-index: 7
Yutong Bian
Yutong Bian
Citations: 6
h-index: 1

Chain-of-Thought (CoT)는 대규모 언어 모델(LLM)의 성능을 향상시키며, 다중 모달 대규모 언어 모델(MLLM)에까지 확장되었습니다. 최근 연구에서는 텍스트 기반의 다중 모드 추론에서 벗어나 중간 단계에서 텍스트 설명과 시각적 증거를 모두 활용하는 통합된 추론 방식으로 발전하고 있습니다. 본 연구에서는 더욱 과감하고 야심찬 아이디어를 제안합니다. 이미지 자체만으로 언어 및 다중 모드 작업에 대한 추론 매개체가 될 수 있을까요? 이를 탐구하기 위해, 우리는 이미지를 독립적인 추론 매체로 취급하는 '광학적 추론'을 제안합니다. 우리는 이 개념을 두 가지 방식으로 구현했습니다. 첫째는 시각적 레이아웃을 최적화하여 간결한 설명을 렌더링하는 '타이포그래피 기반 광학적 추론', 둘째는 텍스트와 그래픽 요소를 구조화된 시각적 설명으로 구성하는 '그래픽 기반 광학적 추론'입니다. 수학, 과학 및 통합 다중 모드 추론 벤치마크에서 광학적 추론은 기존의 텍스트 기반 추론과 동등하거나 그 이상의 성능을 보이며, 언어 작업에서는 평균적으로 28.57%, 다중 모드 작업에서는 16%의 토큰 사용량을 줄여 텍스트 기반 추론보다 1.96배 더 효율적인 결과를 얻었습니다. 이러한 결과는 이미지가 효과적이고 효율적으로 설명을 인코딩할 수 있으며, 추론을 위한 통합된 시각적 환경을 제공할 수 있음을 보여줍니다.

Original Abstract

Chain-of-Thought (CoT) improves the performance of Large Language Models (LLMs) and has been extended to Multimodal Large Language Models (MLLMs). More recent work further moves from text-based multimodal reasoning toward interleaved-modal reasoning, where intermediate steps can incorporate both textual rationales and visual evidence. In this work, we propose a bolder and more ambitious idea: could images alone serve as the reasoning medium for both language and multimodal tasks? To explore this, we propose optical reasoning, which treats images as a standalone reasoning medium. We instantiate this concept with two variants: typographic-based optical reasoning, which optimizes visual layouts for compact rationale rendering, and graphical-based optical reasoning, which composes text and graphical elements into structured visual rationales. Across mathematical, scientific, and interleaved-modal reasoning benchmarks, optical reasoning can match or even exceed traditional text reasoning while reducing reasoning tokens by an average of 28.57% on language tasks and 16% on multimodal tasks, achieving 1.96 times the token efficiency of text reasoning. These results show that images can effectively and efficiently encode rationales while providing a unified visual canvas for reasoning.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!