2608.02139v1 Aug 03, 2026 cs.CL

점진적인 경험 진화를 통한 자기 개선형 대규모 언어 모델

Self-Improving Large Language Models via Progressive Experience Evolution

Fandong Meng
Fandong Meng
Citations: 7,483
h-index: 42
Shijie Ren
Shijie Ren
Citations: 50
h-index: 5
Hao Zhou
Hao Zhou
Citations: 74
h-index: 3
Xiting Wang
Xiting Wang
Citations: 4,488
h-index: 33
Meng Li
Meng Li
Citations: 0
h-index: 0
Yujie Guo
Yujie Guo
Citations: 35
h-index: 3
Yuetan Chen
Yuetan Chen
Citations: 0
h-index: 0
Ziheng Peng
Ziheng Peng
Citations: 7
h-index: 1
Yunhang Yao
Yunhang Yao
Citations: 1
h-index: 1
Yunlong Liang
Yunlong Liang
Citations: 1,440
h-index: 19
Xunlong Wang
Xunlong Wang
Citations: 0
h-index: 0

자기 개선 능력을 갖춘 대규모 언어 모델(LLM)은 효과적인 정책 최적화뿐만 아니라, 일시적인 상호 작용 경험을 지속 가능한 모델 역량으로 변환하는 체계적인 메커니즘을 필요로 합니다. 기존의 자기 개선 방법론은 여전히 단편적인데, 테스트 시나리오 기반 방법은 명시적으로 경험을 추출할 수 있지만 이를 모델 파라미터에 내재화하지 못하고, 학습 시나리오 기반 최적화 방법은 모델 파라미터를 업데이트할 수는 있지만 전이 가능한 경험을 축적하는 명시적인 메커니즘이 부족합니다. 이러한 패러다임 간의 격차를 해소하기 위해서는 아직 충분히 연구되지 않은 중요한 중간 단계인 '경험 증류(experience distillation)'가 필요합니다. 이러한 문제점을 해결하기 위해, 우리는 extbf{SPEE} ( extbf{S}elf- extbf{P}rogressive extbf{E}xperience extbf{E}volution)라는 통일된 사후 학습 프레임워크를 제안합니다. SPEE는 명시적인 경험 진화와 암묵적인 정책 최적화를 순차적으로 수행합니다. 명시적인 경험 진화 단계에서, SPEE는 여러 상호 작용으로부터 수집된 경로를 분석하여 전이 가능한 경험을 추출, 검증 및 점진적으로 발전시키고, 이를 권한 기반 온-폴리시 자기 증류(OPSD)를 통해 정책에 내재화합니다. 암묵적인 정책 최적화 단계에서는, 보상 기반 강화 학습을 활용하여 이러한 내재된 사전 지식을 바탕으로 새로운 해결 전략을 탐색합니다. 경험 진화 단계에서, 지속적으로 발전하는 글로벌 경험 풀은 성공 및 실패 경로 모두에서 얻은 지식을 통합하고, 낮은 유용성을 가진 경험을 제거하며, 개별 경로에 의해 유발되는 사후적 정당화를 완화합니다. 다섯 가지 수학적 추론 벤치마크에 대한 실험 결과, SPEE는 세 가지 모델 규모에서 테스트 시나리오 및 학습 시나리오 기반 자기 개선 방법보다 일관되게 우수한 성능을 보였습니다. 소스 코드는 https://github.com/rrrsj/SPEE 에서 확인할 수 있습니다.

Original Abstract

Large language models (LLMs) capable of self-improvement require not only effective policy optimization, but also a principled mechanism for transforming transient interaction experience into persistent model capabilities. Existing self-improvement paradigms remain fragmented: test-time methods can explicitly extract experience but cannot internalize it into model parameters, whereas training-time optimization methods can update model parameters but lack an explicit mechanism for accumulating transferable experience. Bridging these two paradigms requires a critical intermediate stage that remains underexplored, namely \emph{experience distillation}. To address this gap, we propose \textbf{SPEE} (\textbf{S}elf-\textbf{P}rogressive \textbf{E}xperience \textbf{E}volution), a unified post-training framework that sequentially performs explicit experience evolution followed by implicit policy optimization. During explicit experience evolution, SPEE reflects on trajectories collected from multiple interactions to extract, verify, and progressively evolve transferable experience, which is subsequently internalized into the policy through privilege-guided On-Policy Self-Distillation (OPSD). During implicit policy optimization, reward-driven reinforcement learning leverages these internalized priors to explore novel solution strategies. In the experience evolution stage, a continuously evolving global experience pool consolidates knowledge from both successful and failed trajectories, filters out low-utility experience, and mitigates post-hoc rationalization induced by individual trajectories. Experiments on five mathematical reasoning benchmarks demonstrate that SPEE consistently outperforms both test-time and training-time self-evolution baselines across three model scales. The source code is available at https://github.com/rrrsj/SPEE.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!