2605.26520v1 May 26, 2026 cs.CV

InterSketch: 자체 수정 시각 스케치 및 단계별 보상을 활용한 교차 추론 모델

InterSketch: An Interleaved Reasoning Model with Self-correcting Visual Sketch and Stepwise Reward

Wei Liu
Wei Liu
Citations: 26
h-index: 3
Zhiwei Ning
Zhiwei Ning
Citations: 22
h-index: 2
Lewei Lu
Lewei Lu
Citations: 3
h-index: 1
Jie Yang
Jie Yang
Citations: 22
h-index: 2
J. Ni
J. Ni
Citations: 83
h-index: 4
Hanming Deng
Hanming Deng
Citations: 1,312
h-index: 13
Wenwen Tong
Wenwen Tong
Citations: 1,791
h-index: 7
Xiang Kong
Xiang Kong
Citations: 240
h-index: 5
Shengnan Ma
Shengnan Ma
Citations: 36
h-index: 4
Ziyi Shang
Ziyi Shang
Citations: 0
h-index: 0
Tao Hu
Tao Hu
Citations: 9
h-index: 2
Yong Xien Chng
Yong Xien Chng
Citations: 93
h-index: 4
Jixuan Ying
Jixuan Ying
Citations: 88
h-index: 4
Zehuan Wu
Zehuan Wu
Citations: 75
h-index: 3
Yuan-Lei Zheng
Yuan-Lei Zheng
Citations: 8
h-index: 2

기존의 비전-언어 모델(VLMs)은 다중 단계의 시각적 추론 능력을 보여주었지만, 그 추론 과정은 비교적 피상적이며 텍스트 중심적인 방식으로 이루어져 복잡한 시각적 문제에 대한 적용 가능성이 제한됩니다. 반면, 인간과 유사한 사고는 일반적으로 장기적인 관점에서 시각-텍스트 연계된 사고 과정을 거칩니다. 이러한 간극을 해소하기 위해, 우리는 자체 수정 및 단계별 보상 메커니즘을 통해 VT-CoT(Visual-Textual Chain-of-Thought) 능력을 향상시키는 교차 추론 모델인 InterSketch를 제안합니다. InterSketch는 외부 도구를 사용하여 중간 시각 스케치를 동적으로 생성하고, 이를 텍스트 기반의 추론 과정과 연결하여 장기적인 시각적 이해 과제에 대한 효과적인 인식 및 논리적 추론을 가능하게 합니다. 특히, 초기 단계에서는 고품질의 교차 VT-CoT 데이터셋을 합성하고, 모델이 다중 단계의 교차 추론 및 자체 수정 능력을 갖도록 반성 메커니즘을 포함합니다. 이후 강화 학습(RL) 단계를 통해, 장기적인 추론 과정에서 발생하는 보상 신호의 희소성을 완화하기 위한 단계별 보상 메커니즘을 설계했습니다. 다양한 시각적 추론 벤치마크 실험 결과는 InterSketch가 Gemini-3-Pro와 같은 독점 모델보다 우수한 성능을 보이는 것을 입증합니다.

Original Abstract

While vision-language models (VLMs) have exhibited multi-turn visual reasoning capabilities, their reasoning trajectories remain relatively shallow and are dominated by a text-centric paradigm, limiting their applicability to complex visual challenges. In contrast, human-like thought typically involves long-horizon reasoning with an interleaved visual-textual chain-of-thought (VT-CoT). To bridge this gap, we introduce InterSketch, an interleaved reasoning model to enhance the VT-CoT capability via self-correcting and stepwise reward mechanisms. InterSketch dynamically generates intermediate visual sketches using external tools and interleaves them with textual reasoning, enabling effective perception and logical reasoning over long-horizon visual understanding tasks. Specifically, in the first cold-start stage, we propose a synthesized high-quality interleaved VT-CoT dataset and include a reflection mechanism to enable the model's capability in multi-turn interleaved reasoning and self-correction. In the subsequent reinforcement learning (RL) stage, we design a stepwise reward mechanism to mitigate the sparsity of reward signals inherent in end-only supervision over long-horizon reasoning. Extensive experiments on visual reasoning benchmarks demonstrate the effectiveness of InterSketch, even outperforming proprietary models such as Gemini-3-Pro.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!