2606.17454v1 Jun 16, 2026 cs.AI

에이전트 경로를 통한 모델 행동 분석

Dissecting model behavior through agent trajectories

Anoop Deoras
Anoop Deoras
Citations: 2,990
h-index: 21
Gaurav Gupta
Gaurav Gupta
Citations: 58
h-index: 4
Vatshank Chaturvedi
Vatshank Chaturvedi
Citations: 18
h-index: 1
Jun Huan
Jun Huan
Citations: 58
h-index: 4

인공지능 에이전트의 성능은 단순한 모델링 문제가 아니라 근본적으로 시스템 문제이다. 모델의 고급 기능은 에이전트 하네스를 통해 구현된다. 따라서, 모델의 가정과 하네스의 동작 간의 불일치는 모델의 잠재력을 온전히 에이전트 성능으로 전환하는 것을 방해할 수 있다. 우리는 이러한 불일치를 '의도-실행 격차'로 정의한다. 즉, 모델이 의도한 것과 하네스가 실제로 실행하는 것 사이의 불일치이다. 우리는 이 의도-실행 격차를 최소화하는 것이 도구 및 실행 루프와 같은 다른 하네스 설계 측면만큼 중요하다는 것을 주장한다. 이러한 하네스와 모델 간의 정렬이 미치는 영향을 보여주기 위해, 'Simple Strands Agent' (SSA)라는 간단하고 사용자 정의 가능한 하네스를 개발했다. SSA는 Claude, Gemini, GPT, Grok, Qwen과 같은 다양한 모델 패밀리에 걸쳐 일반화되는 주요 패턴을 찾고, 소수의 모델별 선호도를 반영하는 것을 목표로 한다. 우리는 두 가지 기여를 한다: (i) 다양한 모델 제공업체에서 보고한 SWE-Pro, SWE-Verified 및 Terminal-Bench-2와 같은 인기 있는 에이전트 벤치마크에서의 'pass@1' 성능을 재현하거나 개선하고, (ii) SSA가 생성한 138,000개의 경로를 분석하여, 최첨단 모델에서 상대적으로 균등하게 나타나는 'pass@1' 지표 이상의 정보를 얻는다. 에이전트 경로를 코드 상태 공간으로 표현함으로써, 문제 해결 행동에 대한 모델 수준의 차이를 관찰한다. 편집 빈도, 테스트 활동 및 단계 전환과 같은 세분화된 지표는 개별 모델이 자율적인 문제 해결의 다양한 단계에서 어떻게 노력을 배분하는지 보여준다.

Original Abstract

AI agent performance is not just a modeling problem, it is fundamentally a systems problem. The advanced capabilities of models are realized through agent harnesses. Therefore, a gap between model assumptions and harness behavior can easily prevent the model's full capabilities from translating into agent performance. We formalize this as the `intent-execution' gap: the mismatch between what the model intends and what the harness executes, and vice versa. We argue that minimizing this intent-execution gap is as important as other aspects of harness design such as tools and execution loops. To illustrate the impact of this harness-model alignment, we develop a simple and customizable harness called `Simple Strands Agent' (SSA). SSA aims to find the bulk of common patterns which generalize across different model families (such as Claude, Gemini, GPT, Grok, Qwen), as well as a small number of model-specific preferences. We make two contributions: (i) we $\textbf{reproduce or improve on the pass@1}$ performance reported by diverse model-provider families on popular agentic benchmarks (SWE-Pro, SWE-Verified and Terminal-Bench-2), and (ii) building on an $\textbf{analysis of 138k trajectories generated by SSA}$, we look beyond the $\texttt{pass@1}$ numbers which tend to be relatively even across frontier models. By representing agent trajectories in code state-spaces, we observe model-level differences in problem-solving behavior. Finer-grained metrics such as edit frequency, testing activity, and phase-transitions reveal how individual models allocate effort across different stages of autonomous problem solving.

0 Citations
0 Influential
10.5 Altmetric
52.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!