2607.14004v1 Jul 15, 2026 cs.AI

에이전트 최적화는 누적 효과를 보이는가? Terminal-Bench 2.0을 이용한 지속 학습 평가

Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0

S. Feizi
S. Feizi
Citations: 14,654
h-index: 50
Wenxiao Wang
Wenxiao Wang
Citations: 171
h-index: 5
Priyatham Kattakinda
Priyatham Kattakinda
Citations: 145
h-index: 3

대부분의 에이전트 최적화 방법에서 보고되는 성능 향상은 일회성으로, 즉 에이전트는 고정된 벤치마크에 대해 최적화되고, 그 결과로 얻어진 개선 사항은 해당 방법의 안정적인 특성이었던 것처럼 보고됩니다. 그러나 이는 실제로 배포되는 에이전트에게 중요한 환경을 테스트하지 못합니다. 실제로는 시간이 지남에 따라 새로운 실패와 새로운 작업이 등장함에 따라 최적화가 반복적으로 적용되기 때문입니다. 이 연구에서 우리가 탐구하는 핵심 질문은 다음과 같습니다. 에이전트 최적화를 통해 얻는 성능 향상이 누적되는가? 즉, 에이전트가 한 번 최적화된 후, 새롭게 나타나는 작업에 대해 다시 최적화할 때, 첫 번째 최적화를 통해 얻은 이점을 훼손하지 않고 추가적인 성능 향상을 얻을 수 있는가? 우리는 Terminal-Bench 2.0의 어려운 작업을 활용하여 설계된 두 단계의 지속 학습 평가를 통해 이 질문을 연구합니다. GEPA, Meta Harness, 그리고 RELAI의 검증 가능한 지속 학습 방법(RELAI-VCL)이라는 세 가지 에이전트-하네스 최적화 방식을 동일한 최적화 예산 하에서 비교했습니다. 세 가지 방법 모두 기존의 정적인 단일 단계 환경에서는 기준 에이전트보다 성능 향상을 보였습니다. 그러나 새로운 작업이 도입되면, 이들 방법은 뚜렷하게 달라지는 경향을 보입니다. GEPA로 최적화된 에이전트는 최적화되지 않은 기준 에이전트보다 낮은 성능을 보이는 반면, Meta Harness는 우수한 성능을 보이지만 두 번째 최적화 예산이 주어지더라도 더 이상의 성능 향상을 얻지 못합니다. 반면 RELAI-VCL은 새로운 작업에 대해 긍정적인 전이 성능을 보이면서도, 해당 작업들이 최적화 목표에 통합된 후에도 지속적으로 성능을 향상시켜 모든 평가 단계에서 가장 높은 합격률을 기록했으며, 전체 생애 평균 합격률에서도 가장 높았습니다 (GEPA의 경우 76.4%, Meta Harness의 경우 64.6%, 기준 에이전트의 경우 58.7%). 우리의 주요 발견은 최적화 이점이 오직 회귀 제어가 최적화 루프에 통합되었을 때만 누적된다는 것입니다. 이는 일반화되지 않는 단순한 해결책에 대한 유도 편향을 제공합니다.

Original Abstract

Most reported gains from agent-optimization methods are one-shot: an agent is optimized against a fixed benchmark and the resulting improvement is reported as if it were a stable property of the method. This does not test the setting that matters for deployed agents, where optimization is applied recursively as new failures and new tasks appear over time. The central question this raises is whether optimizer-driven gains compound: after an agent has been optimized once, can it be optimized again on newly arrived tasks without eroding the gains the first round produced? We study this question with a two-phase continual-learning evaluation built from hard tasks in Terminal-Bench 2.0, comparing three approaches to agent-harness optimization (GEPA, Meta Harness, and RELAI's Verifiable Continual Learning, RELAI-VCL) under identical optimization budgets. All three methods improve over the baseline agent in the conventional, static, single-phase setting. However, once new tasks are introduced, the methods diverge sharply: GEPA's optimized agent transfers below the unoptimized baseline, Meta Harness transfers well but fails to improve further once given a second optimization budget, and RELAI-VCL is the only method that both transfers positively to unseen tasks and continues improving after those tasks are folded into the optimization objective, reaching the highest pass rate at every evaluated stage and the highest lifelong average pass rate overall (76.4% vs. 66.0% for GEPA, 64.6% for Meta Harness, and 58.7% for the baseline). Our key observation was that optimization gains compounded only when regression control was built into the optimization loop, providing an inductive bias against shortcut solutions that fail to generalize.

0 Citations
0 Influential
25 Altmetric
125.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!