2606.05922v1 Jun 04, 2026 cs.AI

회고적 하니스 최적화: 경로 탐색을 통한 자기 선호도를 활용하여 LLM 에이전트 성능 향상

Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts

Xianfeng Tang
Xianfeng Tang
Citations: 714
h-index: 11
Jingying Zeng
Jingying Zeng
Citations: 172
h-index: 8
Xiaohua Jia
Xiaohua Jia
Citations: 73
h-index: 4
Shujie Liu
Shujie Liu
Citations: 312
h-index: 8
Wenbo Pan
Wenbo Pan
City University of Hong Kong
Citations: 153
h-index: 5
Xiangyang Zhou
Xiangyang Zhou
Citations: 41
h-index: 2
Yan Lu
Yan Lu
Citations: 2
h-index: 1
Chin-Yew Lin
Chin-Yew Lin
Citations: 1,508
h-index: 7

AI 에이전트는 복잡한 문제를 해결하기 위해 다양한 기술, 도구 및 워크플로우를 통합적으로 사용합니다. 이러한 통합 시스템(하니스)의 지속적인 개선은 새로운 작업에 적응하는 데 필수적입니다. 그러나 기존 최적화 방법은 일반적으로 정확한 검증 데이터 세트를 필요로 하지만, 실제 환경에서는 이러한 레이블링된 데이터를 확보하기 어렵습니다. 이 문제를 해결하기 위해, 우리는 과거의 실행 경로만을 사용하여 에이전트 하니스를 최적화하는 자기 지도 방식인 회고적 하니스 최적화(RHO)를 소개합니다. 구체적으로, RHO는 과거 경로에서 다양한 난이도의 핵심 작업 집합을 선택하고, 이를 병렬로 다시 해결합니다. 에이전트는 이러한 실행 결과를 자체 검증 및 일관성 분석을 통해 평가하고, 후보 하니스 업데이트를 생성한 후, 자체적인 쌍별 선호도를 기준으로 가장 효과적인 업데이트를 선택합니다. 우리는 소프트웨어 엔지니어링, 기술 작업 및 지식 작업을 포함하는 세 가지 다양한 영역에서 RHO를 평가했습니다. 주목할 만한 점은 단일 최적화 단계만으로도 SWE-Bench Pro의 합격률이 외부 평가 없이 59%에서 78%로 향상되었다는 것입니다. 또한, 우리의 분석 결과 RHO가 기존의 실패 요인을 효과적으로 개선한다는 것을 보여줍니다. 그 결과, 최적화된 하니스는 에이전트의 행동 패턴을 변경하고 장기적인 실행 세션 동안 높은 정확도를 유지합니다.

Original Abstract

AI agents rely on a harness of skills, tools, and workflows to solve complex problems. Continually improving this harness is essential for adapting to new tasks. However, existing optimization methods typically require ground-truth validation sets, yet such labeled data is difficult to acquire in practical deployment settings. To address this problem, we introduce Retrospective Harness Optimization (RHO), a self-supervised method that optimizes the agent harness using only past trajectories. Specifically, RHO selects a diverse coreset of challenging tasks from past trajectories and re-solves them in parallel. The agent analyzes these rollouts using self-validation and self-consistency, then generates candidate harness updates and selects the most effective one by its own pairwise self-preference. We evaluate RHO across three diverse domains, spanning software engineering, technical work, and knowledge work. Notably, a single optimization round improves the pass rate on SWE-Bench Pro from 59% to 78% without any external grading. Furthermore, our analysis demonstrates that RHO effectively targets prior failure modes. As a result, the optimized harness alters the agent's behavior patterns and sustains higher accuracy during long-horizon sessions.

1 Citations
0 Influential
5.5 Altmetric
28.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!