2607.21217v1 Jul 23, 2026 cs.AI

ICAE-벤치: 코드 에이전트를 위한 대화형 프로젝트 구축 평가 도구

ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders

Changyi Xiao
Changyi Xiao
Citations: 43
h-index: 3
Yixin Cao
Yixin Cao
Citations: 59
h-index: 4
Chuyu Zhang
Chuyu Zhang
Citations: 96
h-index: 4
Caijun Xu
Caijun Xu
Citations: 4
h-index: 1
Zhongyuan Peng
Zhongyuan Peng
Citations: 13
h-index: 2
Shibo Hong
Shibo Hong
Citations: 68
h-index: 4
Lin Qiu
Lin Qiu
Citations: 78
h-index: 7
Xuezhi Cao
Xuezhi Cao
Citations: 397
h-index: 7
Jiyuan He
Jiyuan He
Citations: 116
h-index: 6
David Lo
David Lo
Citations: 319
h-index: 6
Dan Huang
Dan Huang
Citations: 74
h-index: 2

최근 등장하고 있는 바이브(vibe) 코딩 워크플로우는 코드 에이전트에게 요구되는 역할에 변화를 가져오고 있습니다. 과거에는 명확하게 정의된 지침에 따라 코드를 완성하는 것이 주요 과제였지만, 현재는 에이전트가 계획 수립, 요구사항 명확화, 도구 활용, 디버깅 및 리포지토리 수준의 구축과 같은 다양한 능력을 결합하여 불완전한 제품 의도를 실행 가능한 소프트웨어로 변환할 것으로 기대됩니다. 그러나 기존 벤치마크는 이러한 변화를 충분히 반영하지 못하고 있으며, 여전히 에이전트를 정적인, 완전하게 정의된 작업에 대해 평가하는 데 집중하고 있습니다. 본 논문에서는 대화형 프로젝트 구축 환경에서 코드 에이전트를 평가하기 위한 벤치마크인 ICAE-벤치를 소개합니다. 핵심 아이디어는 모호한 제품 요구사항부터 시작하여 자동화된 사용자 에이전트를 통해 동적인 파라다임을 시뮬레이션하는 것입니다. 이 설정을 현실적이고 평가 가능하게 만들기 위해, ICAE-벤치는 세 가지 주요 설계 요소를 도입했습니다. 첫째, 제약 없는 모호한 요구사항의 애매함을 피하기 위해 각 작업은 실행 가능한 동작을 가진 실제 오픈 소스 리포지토리에서 파생된 명확성을 기반으로 합니다. 둘째, 고품질의 재현 가능한 사용자 시뮬레이션을 보장하기 위해 ICAE-벤치는 사용자 에이전트 데이터를 활용하여 숨겨진 제약을 드러내도록 설계되었으며, 새로운 요구사항을 부여하거나 구현 세부 사항을 노출하지 않습니다. 셋째, 공정한 오픈 엔드 리포지토리 평가를 위해 ICAE-벤치는 표준화된 블랙박스 테스트와 함께 기능 정확성, 의미론 및 API 유사성, 구조적 충실성, 설계 품질 및 상호 작용 품질과 같은 다차원 진단 도구를 사용합니다.

Original Abstract

The recent emergence of vibe-coding workflows is changing what coding agents are expected to do. Instead of merely completing code under fully specified instructions, agents are increasingly expected to transform incomplete product intent into working software by combining various abilities including planning, requirement clarification, tool use, debugging, and repository-level construction. Yet existing benchmarks have not fully caught up with this shift, evaluating agents on static, fully specified tasks. In this paper, we introduce ICAE-Bench, a benchmark for evaluating coding agents under interactive project-building settings. The basic idea is to start from a fuzzy product requirement, simulating the dynamic paradigm with an automated User Agent. To make this setting both realistic and evaluable, ICAE-Bench introduces three key designs. First, to avoid the ambiguity of unconstrained fuzzy requirements, each task derives ambiguity from a precise real open-source repository with executable behavior. Second, to ensure high-quality and reproducible user simulation, ICAE-Bench grounds interaction through User Agent Data, allowing the User Agent to reveal hidden constraints without inventing new requirements or leaking implementation artifacts. Third, to evaluate open-ended repositories fairly, ICAE-Bench uses standardized black-box tests together with multi-dimensional diagnostics, including functional correctness, semantic and API similarity, structural fidelity, design quality, and interaction quality.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!