2607.12227v1 Jul 14, 2026 cs.AI

에이전트의 하네스 진화 평가 방식 재고

Rethinking the Evaluation of Harness Evolution for Agents

Hanna Hajishirzi
Hanna Hajishirzi
Citations: 7,368
h-index: 31
Zhengyu Chen
Zhengyu Chen
Citations: 83
h-index: 5
Yike Wang
Yike Wang
University of California, Berkeley
Citations: 514
h-index: 8
Shakti Senthil
Shakti Senthil
Citations: 0
h-index: 0
Yulia Tsvetkov
Yulia Tsvetkov
Citations: 1,254
h-index: 15
Pradeep Dasigi
Pradeep Dasigi
Citations: 15
h-index: 2
Teng Xiao
Teng Xiao
Citations: 66
h-index: 5
Huaisheng Zhu
Huaisheng Zhu
Citations: 272
h-index: 8
Zhengyu Hu
Zhengyu Hu
Citations: 0
h-index: 0
Yige Yuan
Yige Yuan
Citations: 57
h-index: 3

본 논문에서는 LLM 에이전트를 위한 자동 하네스 진화의 평가 방식을 다시 검토합니다. 기존 하네스 진화 방법은 단위 테스트 케이스를 사용하여 하네스 구성을 탐색하고, 그 결과를 동일한 공개 벤치마크에서 최종 성능으로 보고합니다. 이러한 방식은 두 가지 근본적인 문제를 야기합니다. 첫째, 하네스 진화 자체는 반복적인 탐색 과정이며, 작업 피드백을 통해 후보 하네스를 반복적으로 평가하고 수정합니다. 에이전트의 테스트 시간 확장과 마찬가지로, 그 효과가 개선된 하네스 설계에서 비롯되는 것인지, 아니면 추가적인 탐색만으로 얻어진 것인지 판단하기 위해서는 동일한 피드백 및 추론 예산을 사용한 간단한 작업 수준 탐색 기준선과 비교해야 합니다. 둘째, 탐색 과정과 최종 평가에 동일한 벤치마크를 사용하므로, 보고되는 성능 향상은 해당 특정 작업 세트에 과적합될 위험이 있습니다. 이러한 문제점을 해결하기 위해, 우리는 하네스 진화를 간단한 테스트 시간 확장 및 발견 기준선과 비교하는 광범위한 평가를 수행하고, 또한 발견된 개선 사항이 일반화되는지 확인하기 위해 보류된 작업에서 진화된 하네스를 평가합니다. GPT-5.4와 Claude Opus 4.6을 사용하여 Terminal-Bench 2.1에서 수행한 실험 결과, 자동 하네스 진화는 간단한 테스트 시간 확장 방법보다 일관되게 성능이 우수하지 않으며, 제한적인 일반화 능력을 보였습니다. 이러한 결과는 자동 하네스 진화의 효과에 대한 중요한 질문을 제기하며, 자동 하네스 설계에 대한 공정한 평가 프로토콜 및 벤치마크 개발의 필요성을 강조합니다. 본 논문의 코드는 https://github.com/rethinking-harness-evolution 에서 확인할 수 있습니다.

Original Abstract

We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness configurations and then report final performance on the same public benchmark. This protocol raises two fundamental concerns. First, harness evolution is itself an iterative search procedure that repeatedly evaluates and revises candidate harnesses using task feedback. As in agentic test-time scaling, it should therefore be compared with simple task-level search baselines under matched feedback and inference budgets to determine whether its gains arise from improved harness design or from additional search alone. Second, because the search and the final evaluation share the same benchmark, the reported gains risk overfitting to that specific task set. To address these concerns, we conduct an extensive evaluation comparing harness evolution with simple test-time scaling and discovery baselines under comparable feedback and inference budgets, and also evaluate evolved harnesses on held-out tasks to assess whether the discovered improvements generalize. Experiments on Terminal-Bench 2.1 with GPT-5.4 and Claude Opus 4.6 show that automatic harness evolution does not consistently outperform simple test-time scaling methods and exhibits limited generalization. Our results raise important questions about the effectiveness of automatic harness evolution and highlight the need for fairer evaluation protocols and benchmarks for automatic harness design. Our code is available at https://github.com/rethinking-harness-evolution.

2 Citations
0 Influential
15.5 Altmetric
79.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!