2606.23678v1 Jun 22, 2026 cs.CV

AIR: 코드 통합 다중 모드 대규모 언어 모델에서의 적응적 교차 추론

AIR: Adaptive Interleaved Reasoning with Code in MLLMs

Haibo Qiu
Haibo Qiu
Citations: 177
h-index: 8
Cong Han
Cong Han
Citations: 27
h-index: 2
Xiaohan Lan
Xiaohan Lan
Citations: 19
h-index: 2
Yujie Zhong
Yujie Zhong
Citations: 136
h-index: 6

OpenAI o3에 의해 촉발된 패러다임 전환 이후, 코드를 활용한 교차 추론은 다중 모드 대규모 언어 모델(MLLM)의 성능 향상을 위한 핵심 연구 분야로 자리 잡았습니다. 기존 연구는 주로 시각-인지 작업에서의 도구 사용에 초점을 맞추고 있습니다. 그러나 이러한 접근 방식은 일반적으로 시각적 조작을 위한 미리 정의된 휴리스틱에 의존하며, 시각 연산에만 집중하기 때문에 수치 계산 문제를 해결하는 데 근본적인 한계를 가지고 있습니다. 본 논문에서는 강화 학습 기반 훈련을 통해 코드 통합 복잡한 수치 계산 작업에서 MLLM의 적응적 교차 추론 능력을 향상시킵니다. 이를 위해, 우리는 다음 세 가지 구성 요소로 이루어진 포괄적인 솔루션을 제안합니다: (1) 두 단계로 구성된 초기 데이터 구축 파이프라인, (2) 강화 학습 데이터셋 큐레이션 전략, 그리고 (3) 교차 추론 경로를 위한 그룹 제약 보상 함수를 활용한 적응적 도구 호출 전략. 광범위한 실험 결과는 그룹 제약 보상 함수를 사용한 강화 학습 훈련 후 평가 벤치마크에서 평균적으로 6.1%p의 성능 향상을 보여주며, 특히 교차 추론 샘플에 대한 정확도가 9.9%p 증가하고, 도구 사용 성공률이 95% 이상으로 상승합니다. 데이터 및 코드는 다음 주소에서 확인할 수 있습니다: https://github.com/CongHan0808/AIR.git.

Original Abstract

Following the paradigm shift initiated by OpenAI o3, interleaved reasoning with code to enhance multimodal large language models (MLLMs) has become a pivotal research frontier. The existing literature focuses primarily on tool-use within vision-perception tasks. However, such approaches typically rely on predefined heuristics for visual manipulation and are inherently incapable of addressing numerical computation problems due to their exclusive focus on visual operations. This paper empowers MLLMs with adaptive interleaved reasoning capabilities through extended reinforcement learning training on code-augmented complex numerical computation tasks. To this end, we propose a comprehensive three-component solution consisting of: a two-stage cold-start data construction pipeline, data filtering strategies for RL dataset curation, and an adaptive tool-invocation strategy leveraging a group-constrained reward function for interleaved reasoning trajectories. Extensive experiments demonstrate that after Reinforcement Learning training with the group-constrained reward function, performance improves by an average of 6.1 percentage points (pp) on evaluation benchmarks. Specifically, the accuracy for interleaved reasoning samples increases by 9.9 pp, and the overall success rate of tool-use exceeds 95%. Our data and code are available at: https://github.com/CongHan0808/AIR.git.

0 Citations
0 Influential
24 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!