2608.01918v1 Aug 03, 2026 cs.LG

HarnessCompass: 일반적이고 효과적인 에이전트 하니스를 위한 자동 하니스 진화 가이드

HarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent Harnesses

Zhengyu Chen
Zhengyu Chen
Citations: 83
h-index: 5
Yan Xu
Yan Xu
Citations: 53
h-index: 4
Ruochen Zhou
Ruochen Zhou
Citations: 244
h-index: 5
Dandan Song
Dandan Song
Citations: 39
h-index: 3
Guangyuan Feng
Guangyuan Feng
Citations: 0
h-index: 0
Jun Yang
Jun Yang
Citations: 34
h-index: 3
Luan Zhang
Luan Zhang
Citations: 28
h-index: 3
Yuhang Tian
Yuhang Tian
Citations: 104
h-index: 6
Huipeng Ma
Huipeng Ma
Citations: 16
h-index: 2
Chenhao Li
Chenhao Li
Citations: 7
h-index: 1
Xudong Li
Xudong Li
Citations: 0
h-index: 0
Yizhou Jin
Yizhou Jin
Citations: 33
h-index: 3

하니스(harness) 설계는 대규모 언어 모델(LLM)이 실행 환경 내에서 어떻게 정보를 인식하고, 추론하며, 행동하는지에 영향을 미쳐 에이전트 성능에 중요한 역할을 합니다. 최근 연구에서는 에이전트-환경 상호작용을 통해 하니스를 반복적으로 개선하는 자동 하니스 진화 방법을 제안했습니다. 그러나 기존 방법은 종종 특정 작업에 과적합되고, 궤적 데이터에서 파생된 신호에만 의존하며, 하니스 구성 요소들을 함께 최적화하여 구성 요소 간의 간섭을 초래합니다. 본 논문에서는 제약 조건 기반 진화, 적극적인 피드백 및 구성 요소별 최적화를 활용한 새로운 자동 하니스 진화 프레임워크인 HarnessCompass를 제안합니다. HarnessCompass는 먼저 전역 제약을 적용하여 진화를 제한함으로써 작업에 독립적인 하니스 변경 사항으로 일반화를 넘어선 수정만 수행하도록 합니다. 또한, 에이전트의 하니스 사용에 대한 적극적인 1인칭 피드백을 궤적 데이터 기반 증거와 결합하여 보다 풍부한 진화 신호를 제공합니다. 마지막으로, 다양한 하니스 구성 요소들을 개별적으로 최적화한 후 통합된 하니스로 결합하여 구성 요소 간의 간섭을 줄이면서 구성 요소 간의 시너지 효과를 유지합니다. SWE-bench Verified에서 GPT-5.4를 사용하여 실험한 결과, HarnessCompass는 5번의 진화 반복만으로 Pass@1 정확도를 54%에서 66%로 향상시켜 AHE보다 효율성과 성능 모두에서 우수한 결과를 보였습니다. 또한, 진화된 하니스는 다른 작업과 모델에 효과적으로 적용되어 기존 자동 하니스 진화 방법보다 훨씬 뛰어난 일반화 능력을 보여줍니다.

Original Abstract

Harness design plays a critical role in agent performance by shaping how large language models (LLMs) perceive, reason over, and act within executable environments. Recent work has proposed automatic harness evolution, which iteratively improves the harness from agent--environment interactions. However, existing methods often overfit to the evolution tasks, rely exclusively on trajectory-derived signals, and optimize harness components jointly, causing interference across components. We propose HarnessCompass, a novel automatic harness evolution framework built around constrained evolution, proactive feedback, and component-wise optimization. HarnessCompass first enforces global constraints on evolution, restricting modifications to task-agnostic harness changes that generalize beyond the evolution tasks. It then augments trajectory-derived evidence with proactive first-person feedback from the agent about harness usage, yielding richer signals for evolution. Finally, it decouples the optimization of different harness components before consolidating them into a unified harness, reducing cross-component interference while preserving component synergy. On SWE-bench Verified with GPT-5.4, HarnessCompass improves Pass@1 from 54\% to 66\% in only 5 evolution iterations, outperforming AHE in both effectiveness and evolution efficiency. In addition, the evolved harness transfers effectively to held-out tasks and other models, demonstrating substantially stronger generalization than prior automatic harness evolution methods.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!