2605.29568v1 May 28, 2026 cs.AI

DeepTool: 프로세스 기반 강화 학습을 통한 도구 통합 추론에서의 계층화된 숙고 확장

DeepTool: Scaling Interleaved Deliberation in Tool-Integrated Reasoning via Process-Supervised Reinforcement Learning

Xiao Ding
Xiao Ding
Citations: 1,264
h-index: 17
Bibo Cai
Bibo Cai
Citations: 94
h-index: 6
Kai Xiong
Kai Xiong
Research Center for Social Computing and Information Retrieval
Citations: 406
h-index: 9
Zhouhao Sun
Zhouhao Sun
Citations: 87
h-index: 5
Bing Qin
Bing Qin
Citations: 411
h-index: 12
Ting Liu
Ting Liu
Citations: 1,454
h-index: 15
Yangfan He
Yangfan He
Citations: 149
h-index: 6
Yufei Zhang
Yufei Zhang
Citations: 35
h-index: 3

도구 통합 추론(TIR)은 외부 환경을 활용하여 LLM의 기능을 확장합니다. 그러나 기존 방법들은 전략적 계획 및 자기 수정에 필요한 순차적인 도구 호출 과정에서 충분한 숙고 과정을 제공하지 못합니다. 강화 학습(RL)은 이러한 문제를 완화할 수 있지만, 기존의 TIR 방식은 결과 기반의 희소 보상으로 인해 중간 추론 단계와 도구 호출을 제대로 감독하지 못하는 한계가 있습니다. 이를 해결하기 위해, 우리는 각 단계에서 사고, 행동 및 관찰 과정을 반복하며 숙고를 확장하는 새로운 프레임워크인 DeepTool을 제안합니다. DeepTool에서는 먼저, 적대적 교란을 통합하여 견고성과 자기 수정을 보장하는 확장된 추론 과정을 계층화된 방식으로 변환하는 합성 파이프라인을 도입했습니다. 둘째, GRPO 기반의 프로세스 감독 강화 학습을 설계하여, 행동 중심의 프로세스 보상을 활용하여 중간 단계의 계층화된 사고를 강화하고 각 단계에서 정확한 도구 호출을 유도합니다. 광범위한 실험 결과, DeepTool은 Qwen2.5-7B 모델에 대해 6가지 벤치마크(예: AIME24: 3.2% -> 40.4%, HMMT25: 0.0% -> 28.6%)에서 상당한 성능 향상을 보여주었습니다. 또한, 토큰 비용 효율성 분석 결과는 계층화된 사고의 유용성을 확인했으며, DeepTool이 성능과 토큰 효율성의 최적 균형을 제공한다는 것을 입증합니다.

Original Abstract

Tool-Integrated Reasoning (TIR) extends LLM capabilities by leveraging external environments. However, existing methods lack the deliberation during sequential tool invocation required for strategic planning and self-correction. While RL mitigates this, conventional approaches for Tool-Integrated Reasoning are hindered by sparse outcome-based rewards, failing to supervise intermediate reasoning steps and tool invocations. To address this, we propose DeepTool, a novel framework that scales deliberate thinking within the interleaved process of thinking, action, and observation at each turn. In DeepTool, we first introduce a synthesis pipeline that evolves extended thinking into interleaved trajectories, integrating adversarial perturbations to ensure robustness and self-correction. Secondly, we devise Process-Supervised Reinforcement Learning based on GRPO, which utilizes an Action-Centric Process Reward to reinforce intermediate interleaved thinking and enforce precise tool invocation at every turn. Extensive experiments demonstrate that DeepTool achieves superior performance, boosting Qwen2.5-7B significantly across six benchmarks (e.g., AIME24: 3.2% -> 40.4% and HMMT25: 0.0% -> 28.6%). Furthermore, the token cost-effectiveness analysis confirms the utility of interleaved thinking, demonstrating DeepTool's optimal balance between performance and token efficiency.

0 Citations
0 Influential
8.5 Altmetric
42.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!