2606.16215v1 Jun 15, 2026 cs.CL

PACT: 권한 부여된 추적 공동 학습을 통한 다중 회전 도구 사용 에이전트

PACT: Privileged Trace Co-Training for Multi-Turn Tool-Use Agents

Yingbin Liang
Yingbin Liang
Citations: 121
h-index: 5
Kejing Xia
Kejing Xia
Citations: 3
h-index: 1
Zhenbang Du
Zhenbang Du
Citations: 28
h-index: 2
Xiangchi Yuan
Xiangchi Yuan
Citations: 78
h-index: 4
Qirui Jin
Qirui Jin
Citations: 117
h-index: 4
Wenke Lee
Wenke Lee
Citations: 54
h-index: 4
Shaofeng Zou
Shaofeng Zou
Citations: 26
h-index: 4
Dachuan Shi
Dachuan Shi
Citations: 510
h-index: 9
Jun Luo
Jun Luo
Citations: 176
h-index: 6
Zhiwei Zheng
Zhiwei Zheng
Citations: 6
h-index: 1
Qijia He
Qijia He
Citations: 15
h-index: 2

다중 회전 도구 사용 에이전트는 여러 상호 작용 단계에 걸쳐 추론하고, 도구를 호출하며, 관찰 결과에 적응해야 합니다. 이러한 에이전트에 대한 사후 학습은 프롬프트 기반 추론과 일치하지만 강화 학습은 종종 희소한 보상과 약한 신용 할당 문제로 어려움을 겪습니다. 반면, 전문가의 추적 데이터를 사용한 지도 미세 조정은 밀집된 프로세스 감독을 제공하지만 모델이 고정된 경로에 과도하게 제약될 수 있습니다. 이러한 문제를 해결하기 위해, 우리는 다중 회전 도구 사용 에이전트를 위한 권한 부여된 추적 공동 학습 프레임워크인 PACT를 제안합니다. 핵심 아이디어는 전문가의 추적 데이터를 롤아웃 시 힌트로 사용하는 것이 아니라 학습 시간 최적화 신호로만 활용하는 것입니다. PACT는 롤아웃 생성을 프롬프트 기반으로 유지하고, 두 가지 상호 보완적인 신호를 사용하여 전문가 추적 데이터로부터 최적화를 유도합니다. 첫째, 프롬프트 기반 롤아웃을 전문가 추적 컨텍스트 하에서 평가하는 추적 조건부 강화 학습(RL) 서프록시입니다. 둘째, 추론 전처리 및 도구 호출에 대해 경사 강하율을 조정하여 감독하는 구성 요소 인식 지도 미세 조정(SFT) 손실입니다. PACT는 또한 학습 데이터에만 존재하는 추적 컨텍스트에 대한 과도한 의존성을 줄이기 위해 프롬프트 기반 앵커링을 도입합니다. 또한, 우리는 두 가지 추적 기반 목표를 연결하고 전문가의 추적 데이터가 롤아웃 생성 중에 사용되지 않더라도 최적화를 어떻게 안내할 수 있는지 설명하는 잠재적 추적 관점을 제시합니다. FTRL, BFCL 및 ToolHop에 대한 실험 결과는 PACT가 강력한 지도 미세 조정 및 강화 학습 기반 모델을 지속적으로 능가하며, 이는 다중 회전 도구 사용 학습에서 권한 부여된 추적 공동 학습의 가치를 강조합니다.

Original Abstract

Multi-turn tool-use agents must reason, call tools, and adapt to observations across several interaction turns. Post-training such agents is challenging, as reinforcement learning often suffers from sparse rewards and weak credit assignment despite matching the prompt-only inference setting, while supervised fine-tuning on expert traces provides dense process supervision but can over-constrain the model to fixed trajectories. To tackle this, we propose PACT, a Privileged trAce Co-Training framework for multi-turn tool-use agents. The key idea is to use expert traces only as training-time optimization signals rather than rollout-time hints. PACT keeps rollout generation prompt-only, then uses expert traces to guide optimization through two complementary signals: a trace-conditioned RL surrogate that evaluates prompt-only rollouts under expert-trace context, and a component-aware SFT loss that supervises reasoning prefixes and tool-calls with annealed strength. To reduce over-reliance on the training-only trace context, PACT further introduces a prompt-only anchoring. We also provide a latent-trace view that connects the two trace-based objectives and explains how expert traces can guide optimization without being used during rollout generation. Experiments on FTRL, BFCL, and ToolHop show that PACT consistently improves over strong SFT- and RL-based baselines, highlighting the value of privileged trace co-training for multi-turn tool-use learning.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!