2608.03474v1 Aug 04, 2026 cs.CV

MT-Web2Code: 다중 단계 지역 재구성 및 로컬 수정에 대한 코딩 에이전트 성능 평가

MT-Web2Code: Benchmarking Coding Agents on Multi-Turn Regional Reconstruction and Localized Modification

Xiaocheng Feng
Xiaocheng Feng
Citations: 10,669
h-index: 31
Guanglu Wan
Guanglu Wan
Citations: 28
h-index: 3
Qiming Li
Qiming Li
Citations: 72
h-index: 3
Hao Liu
Hao Liu
Citations: 0
h-index: 0
Songxiang Liu
Songxiang Liu
Citations: 426
h-index: 5
Shujie Hu
Shujie Hu
Citations: 364
h-index: 6

최근 대규모 시각-언어 모델(LVLM)의 발전은 웹 UI 생성 분야에서 놀라운 능력을 보여주었습니다. 그러나 기존 벤치마크는 대부분 처음부터 전체 페이지를 한 번에 생성하는 데 초점을 맞추고 있어, 실제 프론트엔드 개발 워크플로우에서 개발자들이 누락된 영역을 반복적으로 재구성하고 기존 코드베이스 내의 특정 요소를 수정하는 과정을 간과합니다. 이러한 격차를 해소하기 위해, 우리는 다중 단계 수준에서의 지역 재구성과 미세 수준의 로컬 수정에 대한 최초의 멀티모달 코딩 벤치마크인 MT-Web2Code를 소개합니다. 이 벤치마크는 16개의 다양한 분야에 걸쳐 102개의 작업으로 구성되어 있습니다. 비용이 많이 드는 단계별 인간 어노테이션 없이 결정적인 수리 경로를 구축하기 위해, 구조적 및 스타일적 결함을 초기 페이지에 반복적으로 주입하는 확장 가능한 역-손상 경로 엔진을 개발했습니다. 또한, 대상 영역의 충실도와 영향을 받지 않은 콘텐츠 보존 여부를 측정하는 이중 축 평가 프로토콜을 제안합니다. 여기에서 지역 재구성은 5차원 VLM 기반 기준으로 평가하고, 로컬 수정은 결정적인 픽셀 기반 정렬을 통해 평가합니다. 13개의 최첨단 코딩 에이전트에 대한 실험 결과, 현재 에이전트들은 대상 영역을 정확하게 재구성하면서 동시에 영향을 받지 않은 콘텐츠를 유지하는 데 어려움을 겪고 있으며, 로컬 수정에 필요한 세밀한 시각-코드 정렬 능력이 부족하고, 여러 단계에서 오류가 누적되는 경향이 있음을 보여줍니다. 본 연구는 단순한 벤치마킹을 넘어, 반복적인 UI 코딩 에이전트 학습 연구를 촉진할 수 있는 상세한 피드백 신호를 제공하는 결정론적 평가 지표를 제시합니다. 저희의 평가 코드 및 데이터는 곧 공개될 예정입니다.

Original Abstract

Recent advances in Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in web UI generation. However, existing benchmarks predominantly focus on single-turn full-page generation from scratch, overlooking the iterative workflow of real-world frontend engineering, where developers repeatedly reconstruct missing regions and modify localized elements within existing codebases. To bridge this gap, we introduce MT-Web2Code, the first multimodal coding benchmark for multi-turn Macro-Level Regional Reconstruction and Micro-Level Localized Modification, which contains 102 tasks spanning 16 vertical domains. To construct deterministic repair trajectories without costly turn-level human annotation, we develop a scalable Reverse-Corruption Trajectory Engine that iteratively injects structural and stylistic defects into golden pages. We further propose a dual-axis evaluation protocol that measures target-region fidelity and the preservation of unaffected content, where regional reconstruction is assessed by a 5-dimensional VLM-based rubric and localized modification by deterministic pixel-grounded alignment. Experiments on 13 frontier coding agents reveal that current agents struggle to faithfully reconstruct target regions while preserving unaffected content, lack fine-grained visual-code alignment for localized edits, and suffer from error snowballing over multiple turns. Beyond benchmarking, our deterministic evaluation metrics provide fine-grained feedback signals that may facilitate future research on training iterative UI coding agents. Our evaluation code and data will soon be released.

0 Citations
0 Influential
15.5 Altmetric
77.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!