2608.03689v1 Aug 04, 2026 cs.AI

LiveEvalBench: 웹 생성을 위한 개방형 평가 시스템

LiveEvalBench: Toward Open-World Evaluation for Web Generation

Wei Chen
Wei Chen
Citations: 400
h-index: 5
Jun Zhou
Jun Zhou
Citations: 24
h-index: 3
Ying Tang
Ying Tang
Citations: 56
h-index: 4
Lin Yuan
Lin Yuan
Citations: 105
h-index: 5
Zhen Wen
Zhen Wen
Citations: 154
h-index: 6
Xiaolau Zhang
Xiaolau Zhang
Citations: 0
h-index: 0
Yiyao Wang
Yiyao Wang
Citations: 37
h-index: 3
Y. Fu
Y. Fu
Citations: 4
h-index: 1

최근 대규모 언어 모델은 실행 가능한 프런트엔드 프로젝트를 생성하는 능력이 점점 더 향상되고 있지만, 기존의 벤치마크는 여전히 웹 생성을 정적인 평가 문제로 취급합니다. 우리는 프런트엔드 결과물이 정적이지 않고 상호 작용하며, 다양한 구현 방식이 존재하고, 기존 파이프라인으로는 따라잡기 어려울 정도로 빠르게 진화한다는 점을 강조합니다. 이러한 격차를 해소하기 위해, 웹 생성 평가를 에이전트 기반, 적응형, 확장 가능한 프로세스로 재구성하는 자동화 프레임워크인 LiveEvalBench를 제안합니다. LiveEvalBench는 평가를 협업 리뷰 워크플로우로 구현하며, 빌드 엔지니어, 코드 엔지니어, UI 테스터가 프런트엔드 프로젝트의 전체 수명 주기 동안 증거를 수집합니다. 여기에는 배포 및 코드 검사부터 브라우저 기반 상호 작용까지 포함됩니다. 구현 방식의 다양성을 고려하여, LiveEvalBench는 모델 간 비교 가능성을 위한 공통 기준과 각 결과물에 특화된 구현 기반 평가 기준을 결합하는 적응형 프로토콜을 사용합니다. 또한, 이 프레임워크는 파이프라인 재설계 없이 새로운 평가 역할과 평가 차원을 점진적으로 통합할 수 있도록 지원합니다. 다양한 실제 웹 생성 시나리오에서의 실험 결과, LiveEvalBench가 인간 전문가의 판단과 밀접하게 일치하며 최첨단 모델의 웹 생성 능력을 세밀하게 분석하는 데 유용하다는 것을 보여줍니다. 코드 및 관련 정보는 다음 주소에서 확인할 수 있습니다: https://github.com/wyysteelhead/LiveEvalBench

Original Abstract

Large language models are increasingly capable of synthesizing executable frontend projects, yet existing benchmarks still treat web generation as a static evaluation problem. We argue that frontend artifacts demand a different paradigm: they are interactive rather than static, admit diverse yet equally valid implementations, and evolve faster than rigid pipelines can accommodate. To address these gaps, we present LiveEvalBench, an automated framework that reformulates web-generation evaluation as an agentic, adaptive, and extensible process. LiveEvalBench instantiates evaluation as a collaborative review workflow, in which a Build Engineer, a Code Engineer, and a UI Tester collectively gather evidence across the full lifecycle of a frontend project, from deployment and code inspection to browser-based interaction. To handle implementation diversity, an adaptive protocol couples shared rubrics for cross-model comparability with implementation-grounded criteria tailored to each artifact. The framework further supports incremental integration of new evaluator roles and assessment dimensions without pipeline redesign. Experiments across diverse real-world web-generation scenarios show that LiveEvalBench aligns closely with human expert judgment and provides fine-grained insights into frontier models' web generation capabilities. Code is available at https://github.com/wyysteelhead/LiveEvalBench

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!