2607.27409v1 Jul 29, 2026 cs.SE

SWE-NFI: 비기능적 개선을 위한 코딩 에이전트 연구 및 성능 비교

SWE-NFI: Studying and Benchmarking Coding Agents for Non-Functional Improvements

Hengchen Yuan
Hengchen Yuan
Citations: 59
h-index: 3
Xin Wang
Xin Wang
Citations: 6
h-index: 1
Zhenhao Li
Zhenhao Li
Citations: 825
h-index: 15
Pengyu Xue
Pengyu Xue
Citations: 86
h-index: 5
Junkai Chen
Junkai Chen
Citations: 252
h-index: 8
Haonan Zhang
Haonan Zhang
Citations: 48
h-index: 3
Boyuan Chen
Boyuan Chen
Citations: 151
h-index: 7
Zishuo Ding
Zishuo Ding
Citations: 209
h-index: 7
Weiyi Shang
Weiyi Shang
Citations: 32
h-index: 4

코딩 에이전트는 정확성 중심의 평가 지표에서 뛰어난 성과를 보여주었지만, 동작 방식을 유지하면서 소프트웨어 품질을 향상시키는 비기능적 개선(NFI) 능력이 충분히 연구되지 않았습니다. 실제 소프트웨어 개발에서는 개발자들이 관찰 가능한 동작을 변경하지 않고 지속적으로 소프트웨어 품질을 개선하지만, 기존의 평가 지표는 주로 기능적인 정확성을 평가하며 이러한 비기능적 개선을 측정하는 데 한계가 있습니다. 본 논문에서는 코딩 에이전트의 비기능적 개선 능력을 평가하기 위한 벤치마크인 SWE-NFI를 제시합니다. 저희 벤치마크는 오픈 소스 Python 프로젝트에서 가져온 실제 병합 pull request로부터 구성된 188개의 작업으로 이루어져 있습니다. 개발자 중심적인 비기능적 개선 사항을 92가지의 실행 가능한 규칙으로 정의하고, 기능적 정확성 테스트와 규칙 기반의 비기능적 개선 평가를 결합한 종합적인 평가 시스템을 개발했습니다. 최첨단 상용 및 오픈 소스 코딩 에이전트를 평가한 결과, 가장 뛰어난 성능을 보이는 에이전트도 70.0%의 기능적 정확도를 달성했지만, 전체적으로 평가된 모든 에이전트는 인간 개발자에 비해 비기능적 개선 능력에서 부족했습니다. 특히 코드 구조 개선 측면에서 에이전트의 NFI 점수는 0.0에서 1.3 사이인 반면, 인간 기준으로는 1.5로 나타나 격차가 뚜렷합니다. 저희 벤치마크와 연구 결과는 기능적 정확성 외에 코딩 에이전트를 평가하고 발전시키는 데 활용될 수 있는 재현 가능한 기반을 제공합니다.

Original Abstract

Although coding agents have achieved impressive performance on correctness-oriented benchmarks, their ability to make behavior-preserving non-functional improvements (NFIs) remains underexplored. In real-world software development, developers continuously improve software quality without changing observable behavior, yet existing benchmarks primarily evaluate functional correctness and provide limited support for assessing these non-functional improvements. In this paper, we present SWE-NFI, a benchmark for evaluating coding agents on NFIs beyond functional correctness. Our benchmark contains 188 tasks constructed from real merged pull requests in open-source Python projects. We operationalize developer-oriented NFIs into 92 executable rules and develop a comprehensive evaluation suite that combines functional correctness testing with rule-based NFI evaluation. We evaluate state-of-the-art commercial and open-source coding agents. Although the best-performing agent achieves a 70.0\% functional correctness rate, all evaluated agents generally fall short of human developers in overall NFI capability. The gap is particularly evident for structural code improvements, where agents' NFI scores range from 0.0 to 1.3, compared with 1.5 for the human reference. Our benchmark and findings provide a reproducible foundation for evaluating and advancing coding agents beyond functional correctness.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!