DREvo: 재조정된 과거 경험을 활용한 자기 진화 시스템
DREvo: Distilling Recalibrated Historical Experience for Harness Self-Evolution
언어 모델 에이전트의 성능에서 '하네스(harness)'는 매우 중요한 역할을 하며, 고성능 하네스를 구축하는 데에는 상당한 전문 지식이 필요합니다. 따라서 최근 연구에서는 반복적인 제안, 평가 및 개선을 통해 과거 실험 데이터를 활용하는 '하네스 자기 진화'에 대한 관심이 높아지고 있습니다. 하지만 축적된 과거 데이터가 항상 안정적인 탐색 방향 제시로 이어지는 것은 아니며, 진화 과정에서 성능 변동폭이 커서 제한된 자원 하에서 고성능 하네스를 신뢰성 있게 찾는 것이 어렵습니다. 기존의 하네스 자기 진화 방법들이 과거 경험을 활용하는 데 있어 다음과 같은 두 가지 한계점을 가지고 있다고 판단했습니다: (1) 현재 하네스와 관련된 과거 데이터가 여전히 유효한지 동적으로 재평가하지 못하고, (2) 유효한 과거 데이터를 실행 가능한 탐색 방향으로 명시적으로 변환할 수 있는 메커니즘이 부족합니다. 이러한 한계점을 해결하기 위해, 함수 수준의 증거 고정(function-level evidence anchoring), 상태 의존적 증거 재조정(state-dependent evidence recalibration), 그리고 역할 기반의 탐색 목표 추출(role-conditioned search intent distillation)을 통합하여 과거 데이터의 유효성을 판단하고 하네스가 다음에 어떻게 진화해야 하는지 결정하는 새로운 하네스 자기 진화 방법인 'DREvo'를 제안합니다. 제한된 자원 환경에서 DREvo는 더욱 안정적인 진화 경로를 보여주며, 모든 다섯 가지 벤치마크에서 가장 높은 정확도를 달성했습니다. 또한, 도메인 추론 및 에이전트 관련 작업에서 각각 평균 16.2%와 14.2%의 성능 향상을 보였습니다.
Harness plays a critical role in large language model agent performance, and building a high-performing harness requires substantial expert effort. Therefore, recent research has increasingly explored harness self-evolution, which iteratively proposes, evaluates, and improves harnesses using historical trial experience. However, accumulated historical experience does not always translate into stable search guidance, and performance often fluctuates substantially across evolution iterations, making it difficult to reliably discover high-performing harnesses under a limited evolution budget. We identify two limitations in how existing harness self-evolution methods leverage historical experience: (1) Lack of dynamic reassessment of whether historical experience remains valid for the current harness, and (2) Lack of explicit mechanisms for translating valid historical experience into actionable search directions. To address these limitations, we propose a new harness self-evolution method, named DREvo, which integrates function-level evidence anchoring, state-dependent evidence recalibration, and role-conditioned search intent distillation to determine which historical evidence remains valid and where the harness should evolve next. Under limited evolution budgets, DREvo exhibits smoother evolution trajectories, achieves the highest accuracy on all five benchmarks, and delivers average gains of 16.2% and 14.2% over the evaluated baselines on domain reasoning and agentic tasks, respectively.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.