실무자들이 어떻게 소프트웨어 엔지니어링 에이전트를 구축하는가? 혼합 방법론 연구를 통한 통찰
How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study
소프트웨어 엔지니어링(SE) 에이전트, 즉 대규모 코드베이스를 이해하고 제한적인 인간의 개입으로 엔지니어링 작업을 수행할 수 있는 LLM 기반 에이전트는 빠른 발전과 확산 속도를 보이고 있지만, 실무에서 개발자들이 이러한 시스템을 어떻게 구축하는지에 대한 정보는 부족합니다. 기존 연구들은 저장소 분석이나 배포 과정을 살펴보지만, SE 에이전트의 구축 과정 자체를 심층적으로 조사하는 경우는 드뭅니다. 본 논문은 12개 조직의 20명의 실무자와 온라인 설문에 응답한 80명의 실무자를 대상으로 한 반구조화된 인터뷰와 온라인 설문을 통해, SE 에이전트 개발 과정에서 SE 프로세스가 어떻게 변화하고 있으며, 개발자들이 어떤 어려움에 직면하는지를 처음으로 연구합니다. 연구 결과, 구현 비용이 저렴해짐에 따라 병목 현상은 사라지는 것이 아니라 단순히 다른 영역으로 이동한다는 것을 확인했습니다. 기존의 요구사항 분석, 협업, 검토 및 배포와 같은 비코딩 작업이 더욱 중요해지고, 에이전트 출력의 검토 및 평가가 새로운 핵심 과제로 부상했습니다. 본 연구는 7단계 워크플로우와 평가 중심 개발로의 전환을 제시하며, 여기서 평가는 반복적인 개선을 이끌고, 명세는 인간과 에이전트 모두가 참조하는 버전 관리된 산출물이 됩니다. 또한, 팀들이 직면하는 6가지 과제와 함께, 이러한 과제를 해결하기 위해 채택하는 방법들을 분석합니다. 여기에는 신뢰할 수 없는 평가 지표, 코드 발전 속도가 이해도를 따라가지 못하는 '이해 부채', 그리고 제공업체 측 모델 업데이트로 인해 발생하는 행동 변화 등이 포함됩니다.
The rise of Software Engineering (SE) agents, i.e., LLM-based agents that can understand large codebases and carry out engineering tasks with limited human intervention, has been marked by rapid advances and adoption, but little is known about how developers build these systems in practice: existing studies mine repositories or examine deployment, but few investigate how SE agents are constructed. Through semi-structured interviews with 20 practitioners from 12 organizations and an online survey of 80 practitioners, this paper is the first to study how SE processes are changing in the development of SE agents and what challenges developers face. We find that as implementation becomes cheaper, bottlenecks shift rather than disappear: long-standing non-coding work such as requirements, coordination, review, and deployment becomes more visible, while reviewing and evaluating agent output becomes new and central. We characterize a seven-stage workflow and a shift toward evaluation-driven development, in which evaluation steers iteration and specifications become versioned artifacts read by both humans and agents. We further identify six challenges that teams face, together with the practices they adopt to address them, including unreliable evaluation signals, comprehension debt as code outpaces understanding, and behavioral changes introduced by provider-side model updates.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.