2606.17727v1 Jun 16, 2026 cs.AI

LongWebBench: 장기적인 환경에서의 웹페이지 구조 및 기능 생성 평가

LongWebBench: Evaluating Structural and Functional Webpage Generation in Long-Horizon Settings

Jie Tang
Jie Tang
Citations: 45
h-index: 5
Yi Zhao
Yi Zhao
Citations: 1,134
h-index: 12
Z. Yang
Z. Yang
Citations: 330
h-index: 4
Mengpan Chen
Mengpan Chen
Citations: 12
h-index: 1
Mingde Xu
Mingde Xu
Citations: 23
h-index: 4
Shanghui Gong
Shanghui Gong
Citations: 0
h-index: 0
Xijun Liu
Xijun Liu
Citations: 20
h-index: 2
Jibing Gong
Jibing Gong
Citations: 18
h-index: 2

최근의 시각-언어 모델(VLM)은 시각적 입력으로부터 웹페이지를 생성하는 데 상당한 발전을 보이고 있지만, 기존 평가는 주로 짧고 단일 화면이며 정적인 웹페이지에 초점을 맞추었습니다. 본 논문에서는 구조적 측면과 기능적 측면 모두에서 장기적인 웹페이지 생성을 평가하기 위한 벤치마크인 LongWebBench를 소개합니다. LongWebBench는 구조적 정확성 평가를 위한 490개의 실제 긴 웹페이지와, 129개의 웹페이지에 대한 507개의 목표 지향형 상호 작용 작업을 포함하는 기능 평가 데이터셋입니다. 본 연구에서는 장거리의 구조적 일관성을 평가하기 위한 다차원 VLM 기반 메트릭과, 전체적인 기능 검증을 위한 DOM(Document Object Model) 확장 에이전트 기반 파이프라인이라는 두 가지 보완적인 프로토콜을 사용합니다. 또한, 자동 평가 프로토콜에 대한 인간 동의 분석을 통해 그 타당성을 검토했습니다. 최첨단 오픈 소스 및 독점 VLM 모델을 사용하여 단일 이미지와 다중 이미지 환경에서 실험한 결과, 웹페이지 길이가 증가함에 따라 구조적 정확성이 저하되는 반면, 시각적으로 설득력 있는 생성물조차도 실행 가능한 다단계 상호 작용을 지원하지 못하는 경우가 많다는 것을 확인했습니다. 이러한 결과는 시각적 유사성 외에도 실행 가능한 상호 작용을 핵심 기준으로 평가해야 할 필요성을 강조합니다. 본 연구의 코드 및 데이터는 https://github.com/zheny2751-dotcom/LongWebBench 에서 확인할 수 있습니다.

Original Abstract

Recent vision-language models (VLMs) have shown promising progress in generating webpages from visual inputs, yet existing evaluations mainly focus on short, single-screen, and largely static webpages. We introduce LongWebBench, a benchmark for evaluating long-horizon webpage generation from both structural and functional perspectives. LongWebBench contains 490 real-world long webpages for structural fidelity evaluation and 507 goal-oriented interaction tasks over 129 webpages for functional evaluation. It employs two complementary protocols: a multi-dimensional VLM-based metric for assessing long-range structural coherence, and a DOM-augmented agent-based pipeline for end-to-end functional verification. We further examine the automatic evaluation protocols through human agreement analysis. Experiments with state-of-the-art open-source and proprietary VLMs under single-image and multi-image settings reveal that structural fidelity degrades as webpage length increases, while visually plausible generations often fail to support executable multi-step interactions. These results highlight the need to evaluate long webpage generation beyond visual similarity, with executable interaction as a core criterion. Our code and data are available at https://github.com/zheny2751-dotcom/LongWebBench.

0 Citations
0 Influential
25.5 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!