2605.29442v1 May 28, 2026 cs.SE

코딩 에이전트가 사용자에게 실패를 안기는 방식: 20,574건의 실제 사용 사례 분석을 통한 개발자-에이전트 불일치 현상 연구

How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions

Ningzhi Tang
Ningzhi Tang
Citations: 189
h-index: 7
Gelei Xu
Gelei Xu
Citations: 68
h-index: 4
Yiyu Shi
Yiyu Shi
Citations: 67
h-index: 4
Collin McMillan
Collin McMillan
Citations: 156
h-index: 7
Tao Dong
Tao Dong
Citations: 8
h-index: 1
T. Li
T. Li
Citations: 279
h-index: 8
Chaoran Chen
Chaoran Chen
Citations: 237
h-index: 8
Yu Huang
Yu Huang
Citations: 74
h-index: 5

인공지능 코딩 에이전트는 점점 더 소프트웨어 환경 내에서 직접적으로 작동하고 있지만, 기존의 실패 분석은 벤치마크 데이터만을 사용하여 실제 개발자가 경험하는 불일치를 제대로 반영하지 못합니다. 본 연구에서는 1,639개 저장소에서 수집된 20,574건의 코딩 에이전트 사용 사례를 관찰하여 개발자의 반발을 통해 드러나는 불일치 현상을 분석했습니다. 우리는 불일치를 '개발자의 반발'을 통해 가시화되는 오류로 정의하고, 각 사례를 형태, 원인, 비용, 해결 방식의 네 가지 축으로 분류하여 분석했습니다. 연구 결과, 에이전트가 프로젝트를 이해하는 방식, 개발자 의도를 해석하는 방식, 규칙을 따르는 방식, 행동 범위를 제한하는 방식, 코드 구현 및 실행 방식, 그리고 진행 상황 보고 방식 등 7가지 주요 불일치 형태가 반복적으로 나타나는 것을 확인했습니다. 90.5%의 사례에서 불일치는 심각한 시스템 손상을 초래하기보다는 노력과 신뢰에 대한 비용을 발생시키지만, 91.49%의 경우에도 여전히 명시적인 사용자 수정이 필요합니다. 또한, 불일치 패턴은 IDE와 CLI 환경 간에 차이를 보이며, 연속된 세션에서도 지속되고 시간이 지남에 따라 변화합니다. 전체적으로 불일치 비율은 감소하지만, 제약 위반과 부정확한 자체 보고는 그 비중을 늘리고 있습니다. 본 연구의 결과는 코딩 에이전트를 실제 개발자 워크플로우와 일치하도록 설계하고 훈련하며 평가하고 인터페이스를 구축하는 데 도움이 될 것입니다.

Original Abstract

AI coding agents increasingly act directly within software environments, yet existing analyses of their failures rely on benchmark trajectories that miss how developers actually experience misalignment. We present an observational study of 20,574 coding-agent sessions from 1,639 repositories across IDE and CLI workflows. We operationalize misalignment as a breakdown made visible through developer pushback, and annotate each episode along four axes: form, cause, cost, and resolution. We identify seven recurring forms, spanning how agents read projects, interpret developer intent, follow rules, bound their actions, implement and execute code, and report progress. 90.50\% of episodes impose effort and trust costs rather than irreversible system damage, yet 91.49\% of visible resolutions still require explicit user correction. Misalignment patterns also differ across IDE and CLI settings, persist across adjacent sessions, and shift over time: while overall rates decline, constraint violations and inaccurate self-reporting grow in share. Our findings inform the design of training, evaluation, and interfaces for keeping coding agents aligned with real developer workflows.

5 Citations
0 Influential
4 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!