2607.21419v1 Jul 23, 2026 cs.AI

PATS: 에이전트 강화 학습을 위한 정책 기반 학습 지원 체계

PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning

Zhengzhou Zhu
Zhengzhou Zhu
Citations: 42
h-index: 4
Qitai Tan
Qitai Tan
Citations: 8
h-index: 2
Peng Chen
Peng Chen
Citations: 61
h-index: 3
Yipeng Shi
Yipeng Shi
Citations: 0
h-index: 0
Zhipeng Ma
Zhipeng Ma
Citations: 22
h-index: 2
Yang Li
Yang Li
Citations: 8
h-index: 2
Yue Wang
Yue Wang
Citations: 0
h-index: 0

장기적인 LLM 에이전트 강화 학습에서, 비효율적인 정책은 종종 유사한 실패를 반복하며, 유용한 정보를 제공하지 못하는 실행 경로를 생성하고 효과적인 정책 최적화를 제한합니다. 기존의 기술 중심 방법들은 재사용 가능한 기술을 최적화하거나, 필터링하거나, 내재화하여 탐색을 개선합니다. 하지만 이러한 방법들은 여전히 기술 자체에 초점을 맞추고 있으며, 변화하는 정책에 대한 적응적인 학습 시간 지원으로 설계되지 않았습니다. 이를 해결하기 위해, 우리는 정책 중심의 학습 패러다임을 제안하며, 기술을 정적인 요소가 아닌 동적인 학습 지원체계로 재정의합니다. 우리의 프레임워크인 Pats는 최신 정책에서 생성된 실행 그룹들을 증거 카드로 변환하고, 작업별 평가를 사용하여 이후 실행에 사용되는 컨텍스트를 조정합니다. 구체적인 지침은 비효율적인 정책이 어려운 작업을 완료하도록 돕습니다. 정책이 개선됨에 따라, 불필요한 컨텍스트는 수정되거나 제거되어 명시적인 지침에 대한 의존성을 줄이면서 유용한 실행 경로의 다양성을 유지합니다. 정책은 표준 RLVR을 사용하여 환경 보상을 통해 최적화되며, 학습 지원체계는 배포 시 삭제됩니다. ALFWorld 및 WebShop에서 Pats는 강력한 기준 모델보다 최대 18.6% 향상된 성능을 보였습니다. 또한, 일곱 가지 검색 증강 질의응답 벤치마크에서 Pats는 32.1% 더 적은 프롬프트 토큰을 사용하면서도 경쟁력 있는 성능을 유지했습니다.

Original Abstract

In long-horizon LLM agent reinforcement learning, weak policies often repeat similar failures, producing uninformative rollout trajectories and limiting effective policy optimization. Existing skill-centric methods improve exploration by optimizing, filtering, or internalizing reusable skills. However, they remain centered on the skills themselves rather than being designed as adaptive training-time support for the evolving policy. To address this, we propose a policy-centric training paradigm that reframes skills as a dynamic training scaffold. Our framework, Pats, converts rollout groups from the latest policy into evidence cards and uses task-specific evaluation to adjust the context used in subsequent rollouts. Concrete guidance helps weak policies to complete challenging tasks. As policy improves, redundant context is revised or removed to reduce reliance on explicit guidance while preserving useful rollout variation. The policy is optimized with environmental rewards using standard RLVR, and the training scaffold is discarded at deployment. On ALFWorld and WebShop, Pats improves over strong baselines by up to 18.6%. Across seven search-augmented QA benchmarks, it remains competitive while using 32.1% fewer prompt tokens than the baseline.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!