2607.24720v1 Jul 27, 2026 cs.CL

다중 단계 장기 계획의 물리학적 원리: 사전 학습부터 사후 학습까지, 단일 및 다중 지도 기반 온정책 에이전트 지식 증류를 통한 연구

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

Zhuoran Jin
Zhuoran Jin
Citations: 607
h-index: 13
Kang Liu
Kang Liu
Citations: 220
h-index: 9
Tianyi Men
Tianyi Men
Citations: 104
h-index: 5
Jun Zhao
Jun Zhao
Citations: 901
h-index: 17

멀티턴(multi-turn) 방식의 장기 계획은 기초 모델 에이전트에게 매우 중요하지만, 어떻게 근본적으로 개선할 수 있는지는 아직 명확하지 않습니다. 기존 모델들은 통제 불가능하고 투명성이 낮은 인터넷 데이터를 기반으로 학습되기 때문에, 계획 능력이 어떻게 습득되고 형성되며 통합되는지 파악하기 어렵습니다. 이러한 문제를 해결하기 위해, 우리는 정밀한 제어가 가능한 통합적이고 통제된 멀티턴 환경을 도입했습니다. 이를 통해 세 단계에 걸쳐 장기 계획 과정을 체계적으로 연구할 수 있습니다. (1) 사전 학습 과정에서의 계획 능력 습득: 데이터 형식, 분포 및 품질을 분석합니다. Chain-of-Thought(CoT)를 활용한 명시적인 세계 모델 구축은 더 강력한 장기 일반화 성능을 제공합니다. 개별적인 기술만으로는 합성적 일반화를 달성하기 어렵지만, 적절한 양의 장기 데이터를 활용하면 가능합니다. 또한, 최적이 아닌 경로들은 오류가 장기적으로 증폭되어 성능 저하를 야기합니다. (2) GRPO 및 OPD 사후 학습을 통한 계획 능력 형성: 상호 정보(mutual information)를 통해 일반적인 계획 패턴과 특정 작업에 관련된 지식을 구분합니다. 계획 패턴의 경우, 사후 학습의 세 가지 적용 영역(불필요, 효과적, 지원 불가)을 식별했습니다. 낮은 품질 및 장기 환경에서 OPD는 GRPO보다 더 넓은 효과적인 영역을 갖습니다. 이는 OPD가 보다 일관된 업데이트 방향을 제공하기 때문입니다. 계획 지식의 경우, 서로 다른 지식을 가진 지도 모델로부터 새로운 절차를 학습시키는 과정은 학생 모델의 기존 세계 모델링 능력을 저해할 수 있으며, 새로운 지식이 완전히 확립되지 않을 수도 있습니다. (3) MOPD 사후 학습을 통한 계획 능력 통합: 멀티-티처(multi-teacher) 온정책 지식 증류(MOPD)는 다양한 환경에서 공유되는 계획 패턴으로 수렴함으로써 에이전트의 능력을 통합할 수 있음을 보여줍니다. 호환 가능한 패턴은 환경 간 일반화를 가능하게 하고, 부분적으로 공유되는 패턴은 지속적인 학습을 지원하며, 완전히 충돌하는 패턴은 심각한 간섭을 초래합니다.

Original Abstract

Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are trained on uncontrollable and opaque Internet data, making it difficult to identify how planning ability is acquired, shaped, and integrated. To address this challenge, we introduce a unified and controlled multi-turn environment that enables precise control. It allows systematically study long-horizon planning across three stages. (1) Planning ability acquisition during pre-training. We study data format, distribution, and quality. Explicit world model construction through CoT state transition modeling yields stronger long-horizon generalization. Atomic skills alone are insufficient for compositional generalization, whereas a litte long-horizon data works. Moreover, suboptimal trajectories severely impair performance because errors amplify over long horizons. (2) Planning ability shaping via GRPO and OPD post-training. Through mutual information, we distinguish general planning patterns from task-specific planning knowledge. For planning patterns, we identify three application regions of post-training: unnecessary, effective, and unsupported. OPD has a broader effective region than GRPO under low-quality and long-horizon settings, as it provides more consistent update directions. For planning knowledge, distilling unseen procedures from a teacher with different knowledge may impair student's prior world modeling without fully establishing new knowledge. (3) Planning ability integration through MOPD post-training. We show that multi-teacher on-policy distillation (MOPD) integrates capabilities by converging to shared planning-pattern across environments. Compatible patterns enable cross-environment generalization, partially shared patterns support continual learning, while completely conflicting patterns cause severe interference.

2 Citations
0 Influential
8.5 Altmetric
44.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!