Harness-Bench: 현실적인 에이전트 워크플로우에서 모델 간의 하니스 효과 측정
Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows
LLM 에이전트는 점점 더 도구를 사용하고, 작업 공간을 수정하며, 구체적인 결과물을 생성하는 실행 가능한 시스템으로 활용되고 있습니다. 이러한 워크플로우에서 성능은 기본 모델뿐만 아니라 '하니스'라고 하는 시스템 계층에 의해서도 크게 영향을 받습니다. 하니스는 컨텍스트, 도구, 상태, 제약 조건, 권한, 추적 및 복구를 관리합니다. 그러나 기존 벤치마크는 일반적으로 실행 과정을 단순화하거나, 전체 에이전트 시스템을 비교하거나, 하니스를 고정하여 실행 계층의 변동성을 연구하기 어렵게 만듭니다. 본 논문에서는 현실적인 에이전트 워크플로우에서 구성 수준의 하니스 효과를 평가하는 진단 벤치마크인 Harness-Bench를 소개합니다. Harness-Bench는 다양한 모델 백엔드에 대해 공통된 작업 환경, 예산 및 평가 프로토콜을 사용하면서 각 하니스의 고유한 실행 동작을 유지하며, 대표적인 하니스 구성을 평가합니다. 이 벤치마크에는 실제 에이전트 활용 패턴에서 파생되어 현실성, 해결 가능성, 올바른 답 검증 가능성 및 무결성이 수동으로 검토된 106개의 격리된 오프라인 작업이 포함되어 있습니다. 각 실행 결과는 최종 결과물, 실행 추적 정보, 사용 통계 및 검증 결과를 기록하여 단순히 완료 여부를 넘어 분석할 수 있도록 합니다. 5,194번의 실행 과정을 통해 모델과 하니스의 조합에 따라 완료율, 과정 품질, 효율성 및 오류 발생 패턴에서 상당한 차이가 있음을 확인했습니다. 이러한 결과는 에이전트의 성능을 보고할 때 기본 모델뿐만 아니라 모델-하니스 구성 수준으로 보고해야 함을 시사합니다. 또한 분석 결과, 논리적인 추론과 도구 피드백, 작업 공간 상태, 증거 또는 검증 가능한 출력 계약 간에 연결이 끊어지는 실행 관련 오류가 반복적으로 발생하는 것을 확인했습니다. Harness-Bench는 신뢰성 있고 효율적이며 감사 가능한 에이전트 실행 시스템을 진단하고 개선할 수 있는 재현 가능한 기반을 제공합니다.
LLM agents are increasingly deployed as executable systems that use tools, modify workspaces, and produce concrete artifacts. In such workflows, performance depends not only on the base model, but also on the harness: the system layer that manages context, tools, state, constraints, permissions, tracing, and recovery. However, existing benchmarks typically abstract away execution, compare complete agent systems, or hold the harness fixed, making execution-layer variation difficult to study. We introduce Harness-Bench, a diagnostic benchmark for evaluating configuration-level harness effects in realistic agent workflows. Harness-Bench evaluates representative harness configurations across multiple model backends under shared task environments, budgets, and evaluation protocols, while preserving each harness's native execution behavior. The benchmark contains 106 sandboxed offline tasks constructed from practical agent-use patterns and manually reviewed for realism, solvability, oracle-checkability, and integrity. Each run records final artifacts, execution traces, usage statistics, and validator outputs, enabling analysis beyond final completion. Across 5,194 execution trajectories, we observe substantial variation in completion, process quality, efficiency, and failure behavior across model-harness pairings. These results suggest that agent capability should be reported at the model-harness configuration level rather than attributed to the base model alone. Our analysis further identifies recurring execution-alignment failures, where plausible reasoning becomes decoupled from tool feedback, workspace state, evidence, or verifiable output contracts. Harness-Bench provides a reproducible foundation for diagnosing and improving reliable, efficient, and auditable agent execution stacks.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.