Cookie-Bench: 웹 생성을 위한 지속적인 화면 기반 키 상호 작용 평가
Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation
최첨단 LLM(Large Language Model)의 출시와 함께 프론트엔드 웹 코드는 핵심 제품으로 자리 잡고 있지만, 개발 속도에 맞춰 이러한 인터랙티브 애플리케이션을 평가하는 것은 여전히 비용이 많이 듭니다. Arena와 같은 인간 심사 기반 순위 시스템은 확장성이 부족하기 때문입니다. 기존의 자동화된 프록시는 일반적으로 참조 구현, 테스트 스위트 또는 엄격한 체크리스트에 의존하며, 실제 세션에서 인간 검토자가 수행하는 논리적인 분석을 놓치는 경향이 있습니다. 본 연구에서는 참조 없이도, 자율적으로 작동하고, 종합적인 추론을 가능하게 하는 새로운 평가 체계를 제시하고, 이를 두 가지 아티팩트를 통해 구현합니다. \textbf{\dataname}은 정적 프레젠테이션 및 인터랙티브 애플리케이션 작업을 모두 포함하는 11개 도메인, 54개의 세부 항목, 1,000개의 쿼리로 구성된 WebDev 벤치마크입니다. 이 벤치마크는 세 가지 난이도 수준과 세 가지 대상 언어 그룹으로 균형을 이루며, 유출된 프롬프트로부터 답변을 떠올리는 것을 방지하기 위해 질문 내용이 재구성되었습니다. \textbf{\framename}은 Flavell의 메타인지 모니터링 개념에 기반하여 증거 수집과 판단 과정을 세 단계로 분리합니다. 첫째, Static Perception은 수동 관찰을 통해 초기 인상을 형성합니다. 둘째, Agent-Driven Interaction은 애플리케이션이 자율적으로 작동하는 동안 지속적인 화면 비디오, 오디오 및 각 단계별 스크린샷을 캡처합니다. 셋째, Dynamic Scoring은 증거 체인이 완료된 후에 전체 기능과 미적 측면에 대한 종합적인 판단을 내리고, 구조화된 오류 원인 분석을 제공합니다. \textbf{\framename}은 \textbf{\dataname}에서 전문가의 인간 평가와 밀접하게 일치하며, 13개의 최첨단 LLM이 인터랙티브 웹 생성 분야에서 상당한 잠재력을 가지고 있음을 보여줍니다. 자세한 내용은 다음 링크에서 확인하실 수 있습니다: https://anonymous.4open.science/r/Cookie-3CE/
Front-end web code has become a core product surface for every frontier LLM release, yet evaluating these interactive applications at development speed remains costly because human-judged leaderboards like Arena do not scale. Existing automated proxies typically lean on reference implementations, test suites, or rigid checklists, and tend to miss the reasoned synthesis a human reviewer performs over a live session. We articulate a new evaluation regime that is simultaneously reference-free, autonomously driven, and holistically reasoned, and instantiate it through two artifacts. \textbf{\dataname} is an 11-domain, 54-leaf, 1,000-query WebDev benchmark spanning both static-presentation and interactive-application tasks, balanced across three difficulty tiers and three target-language groups, with briefs rewritten to resist recall from circulated prompts. \textbf{\framename}, grounded in Flavell's metacognitive monitoring, separates evidence accumulation from judgment across three stages: Static Perception forms a first impression from passive observation; Agent-Driven Interaction explores the application autonomously while capturing continuous screen video, audio, and per-step screenshots; Dynamic Scoring issues holistic functionality and aesthetics verdicts with structured failure attribution only after the evidence chain is complete. On \dataname, \framename aligns closely with expert human ratings while surfacing substantial headroom across 13 frontier LLMs on interactive web generation. \noindenthttps://anonymous.4open.science/r/Cookie-3CE/
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.