2607.06306v1 Jul 07, 2026 cs.SE

UI2App: 실행 가능한 웹 애플리케이션 생성 시 시각적 상호 작용 추론 성능 평가

UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation

Yuyu Luo
Yuyu Luo
Citations: 269
h-index: 9
Yifan Wu
Yifan Wu
Citations: 158
h-index: 8
Yiyu Chen
Yiyu Chen
Citations: 12
h-index: 2
Yenchi Tseng
Yenchi Tseng
Citations: 8
h-index: 1
Ying-Cong Chen
Ying-Cong Chen
Citations: 31
h-index: 3
Grace Man Chen
Grace Man Chen
Citations: 0
h-index: 0
Litao Guo
Litao Guo
Citations: 29
h-index: 4
Sicheng Liu
Sicheng Liu
Citations: 0
h-index: 0

대규모 언어 모델(LLM)은 웹 페이지 생성 능력에서 상당한 발전을 보여왔습니다. 그러나 기존의 텍스트 기반 접근 방식은 복잡한 프롬프트를 필요로 하며, 이는 사용자에게 큰 부담을 주고 페이지 레이아웃 및 여러 페이지 간의 시각적 일관성을 표현하는 데 한계가 있습니다. UI 스크린샷을 입력으로 사용하는 이미지 기반 패러다임은 실제 개발 워크플로우와 더 밀접하게 관련되어 있습니다. 그러나 현재 벤치마크는 주로 생성된 결과물의 시각적 충실도에 초점을 맞추고 있으며, 생성된 결과물의 상호 작용 기능에 대한 체계적인 평가가 부족합니다. 이러한 격차를 해소하기 위해, 우리는 상호 작용 추론을 목표로 하는 최초의 벤치마크인 UI2App을 소개합니다. UI2App은 텍스트 또는 행동 지침 없이 스크린샷만으로 애플리케이션 동작을 복원하는 능력, 즉 상호 작용 추론 능력을 평가합니다. UI2App은 실행 가능한 다중 경로 웹 애플리케이션을 구성하는 45개의 상태 일관성 스크린샷 세트로 총 327개의 스크린샷으로 구성됩니다. 우리는 각 결과물을 실행 가능성, 탐색 접근성, 시각적 충실도 및 상호 작용 추론의 네 가지 측면에서 평가하는 엔드 투 엔드 파이프라인을 설계했습니다. 상호 작용 지표(IIS)는 기능적 정확성과 상태 관리 복잡성을 기준으로 추론된 상호 작용을 평가하며, 단일 참조 값과 일치하는 것이 아니라 유효한 구현이면 점수를 부여합니다. 6개의 최첨단 비전-언어 모델에 대한 실험 결과, 시각적 재구성 능력과 실제 상호 작용 구현 능력 사이에 상당한 격차가 있음을 보여줍니다. 시각적 충실도 측면에서 가장 높은 성능을 보이는 모델은 IIS에서 7.5점을 기록하며, 이는 네 번째 순위이며 IIS 최고 점수 모델보다 5.2배 낮은 수치입니다. 여러 페이지에 걸친 상태 관리와 같은 복잡한 상호 작용은 여전히 주요 장애물이며, 평가된 모델의 절반이 이 측면에서 0점을 받았습니다. 전반적으로, 이러한 결과는 정적인 스크린샷으로부터 완전한 상호 작용 동작을 추론하는 것이 모델에게 있어 중요한 과제임을 시사합니다.

Original Abstract

Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing text-driven approaches rely on complex prompts that impose substantial demands on users and offer limited expressivity for page layout and cross-page visual coherence. Image-driven paradigms, which take UI screenshots as input, align more closely with real development workflows. However, current benchmarks focus primarily on visual fidelity and lack a systematic evaluation of the interaction capabilities in generated artifacts. To address this gap, we introduce UI2App, the first benchmark targeting interaction inference, the ability to recover application behavior from screenshots alone, without any textual or behavioral guidance. UI2App comprises 327 screenshots grouped into 45 state-coherent screenshot sets for runnable multi-route web applications. We design an end-to-end pipeline that evaluates each artifact along four dimensions: executability, navigation reachability, visual fidelity, and interaction inference. The interaction metric (IIS) assesses inferred interactions by functional correctness and state-management complexity, crediting any valid implementation rather than matching a single reference. Experiments on six frontier vision-language models reveal a marked capability mismatch between visual reconstruction and interaction realization: the visual-fidelity leader scores only 7.5 on IIS, ranking fourth and trailing the IIS leader by 5.2x. High-complexity interactions such as cross-page state remain a pervasive bottleneck, with half of the evaluated models scoring exactly zero on this dimension. Overall, the results indicate that inferring complete interaction behavior from static screenshots remains a key challenge for models.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!