InterSketch: 자체 수정 시각 스케치 및 단계별 보상을 활용한 교차 추론 모델
InterSketch: An Interleaved Reasoning Model with Self-correcting Visual Sketch and Stepwise Reward
기존의 비전-언어 모델(VLMs)은 다중 단계의 시각적 추론 능력을 보여주었지만, 그 추론 과정은 비교적 피상적이며 텍스트 중심적인 방식으로 이루어져 복잡한 시각적 문제에 대한 적용 가능성이 제한됩니다. 반면, 인간과 유사한 사고는 일반적으로 장기적인 관점에서 시각-텍스트 연계된 사고 과정을 거칩니다. 이러한 간극을 해소하기 위해, 우리는 자체 수정 및 단계별 보상 메커니즘을 통해 VT-CoT(Visual-Textual Chain-of-Thought) 능력을 향상시키는 교차 추론 모델인 InterSketch를 제안합니다. InterSketch는 외부 도구를 사용하여 중간 시각 스케치를 동적으로 생성하고, 이를 텍스트 기반의 추론 과정과 연결하여 장기적인 시각적 이해 과제에 대한 효과적인 인식 및 논리적 추론을 가능하게 합니다. 특히, 초기 단계에서는 고품질의 교차 VT-CoT 데이터셋을 합성하고, 모델이 다중 단계의 교차 추론 및 자체 수정 능력을 갖도록 반성 메커니즘을 포함합니다. 이후 강화 학습(RL) 단계를 통해, 장기적인 추론 과정에서 발생하는 보상 신호의 희소성을 완화하기 위한 단계별 보상 메커니즘을 설계했습니다. 다양한 시각적 추론 벤치마크 실험 결과는 InterSketch가 Gemini-3-Pro와 같은 독점 모델보다 우수한 성능을 보이는 것을 입증합니다.
While vision-language models (VLMs) have exhibited multi-turn visual reasoning capabilities, their reasoning trajectories remain relatively shallow and are dominated by a text-centric paradigm, limiting their applicability to complex visual challenges. In contrast, human-like thought typically involves long-horizon reasoning with an interleaved visual-textual chain-of-thought (VT-CoT). To bridge this gap, we introduce InterSketch, an interleaved reasoning model to enhance the VT-CoT capability via self-correcting and stepwise reward mechanisms. InterSketch dynamically generates intermediate visual sketches using external tools and interleaves them with textual reasoning, enabling effective perception and logical reasoning over long-horizon visual understanding tasks. Specifically, in the first cold-start stage, we propose a synthesized high-quality interleaved VT-CoT dataset and include a reflection mechanism to enable the model's capability in multi-turn interleaved reasoning and self-correction. In the subsequent reinforcement learning (RL) stage, we design a stepwise reward mechanism to mitigate the sparsity of reward signals inherent in end-only supervision over long-horizon reasoning. Extensive experiments on visual reasoning benchmarks demonstrate the effectiveness of InterSketch, even outperforming proprietary models such as Gemini-3-Pro.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.