2608.04007v1 Aug 04, 2026 cs.CL

TurnSight: 도구 통합 추론을 위한 턴 레벨 후향적 자기 증류

TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning

Hengyi Cai
Hengyi Cai
Citations: 531
h-index: 10
Sunhao Dai
Sunhao Dai
Citations: 56
h-index: 4
Yuqi Zhou
Yuqi Zhou
Citations: 237
h-index: 8
Changle Qu
Changle Qu
Citations: 395
h-index: 5
Jun Xu
Jun Xu
Citations: 452
h-index: 7
Xinran Chen
Xinran Chen
Citations: 740
h-index: 12
Simon
Simon
Citations: 108
h-index: 1

도구 통합 추론(TIR)은 LLM이 반복적인 도구 상호 작용을 통해 복잡한 작업을 해결할 수 있도록 합니다. 그러나 기존 강화 학습 방법은 종종 경로 수준의 감독 신호에 의존하여, 장기적인 TIR 시나리오에서 세밀한 보상 할당을 제한합니다. 온-정책 자기 증류는 특권적 컨텍스트를 가진 교사 네트워크를 통해 더 밀집된 신호를 제공하지만, 기존 접근 방식은 일반적으로 이러한 컨텍스트를 정답 또는 검색된 기술로부터 파생시키는데, 이는 에이전트가 실제로 방문하는 상태를 반영하지 못할 수 있습니다. 또한 토큰 수준의 감독은 도구 상호 작용의 턴 레벨 구조를 포착하지 못합니다. 이를 해결하기 위해, 우리는 실행 조건에 따른 후향적 정보를 직접 활용하여 턴 레벨의 자체 신호를 생성하는 TurnSight라는 새로운 프레임워크를 제안합니다. TurnSight는 다양한 예측 범위를 가진 여러 개의 후향적 시점을 구성하고, 서로 다른 시간 간격에서의 방향 일치를 통해 신뢰할 수 있는 감독 신호를 선택합니다. 마지막으로, 선택된 후향적 신호는 관련된 모든 실행 경로에 대해 정규화되고 RL 이점을 적응적으로 조절하는 데 사용되며, 원래의 최적화 방향은 유지됩니다. 세 가지 벤치마크에서의 광범위한 실험 결과는 TurnSight의 효과를 입증합니다. 우리의 코드는 https://github.com/quchangle1/TurnSight 에서 확인할 수 있습니다.

Original Abstract

Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, existing reinforcement learning methods often rely on trajectory-level supervision, limiting fine-grained credit assignment in long-horizon TIR scenarios. On-policy self-distillation offers denser signals through teacher branches with privileged context, but existing approaches typically derive such context from ground-truth answers or retrieved skills, which may not reflect the states actually visited by the agent. Moreover, token-level supervision fails to capture the turn-level structure of tool interactions. To address this, we propose TurnSight, a turn-level hindsight self-distillation framework that derives supervision directly from execution-conditioned hindsight. It then constructs multiple hindsight views with different lookahead horizons and selects reliable supervision through cross-horizon directional agreement. Finally, the selected hindsight signal is normalized across sibling rollouts and used to adaptively modulate RL advantages while preserving their original optimization direction. Extensive experiments on three benchmarks demonstrate the effectiveness of TurnSight. Our codes are available at https://github.com/quchangle1/TurnSight.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!