2608.02603v1 Aug 03, 2026 cs.CV

WorldExam: 겉보기 이미지부터 근본적인 반응성까지, 월드 모델의 성능 평가

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Tieniu Tan
Tieniu Tan
Citations: 146
h-index: 3
Liang Tan
Liang Tan
Citations: 0
h-index: 0
Junhan Zeng
Junhan Zeng
Citations: 0
h-index: 0

제어 가능한 비디오 생성 모델은 점점 더 많은 분야에서 월드 모델로 활용되고 있습니다. 따라서 이러한 모델을 평가할 때는 단순히 생성된 비디오의 시각적 품질뿐만 아니라, 모델이 묘사하는 세계의 고유한 반응성, 즉 장면 상태를 기반으로 세계가 어떻게 반응해야 하는지 추론하고 입력에 명시적으로 설명되지 않은 타당한 결과를 생성하는 능력까지 고려해야 합니다. 하지만 기존의 평가 지표들은 주로 시각적 품질이나 명시적인 명령어 수행 여부를 확인하여 평가하며, 월드 모델의 고유한 반응성을 충분히 검증하지 못합니다. 본 연구에서는 시각적 품질, 제어 준수성, 공간 일관성 및 월드 반응성의 네 가지 수준으로 구성된 계층적 진단 벤치마크인 WorldExam을 소개합니다. WorldExam은 총 1,474개의 사례를 포함하며, 8가지의 특정 작업에 적용 가능하도록 설계되어 있으며, 카메라, 액션, 언어 기반 모델의 통합적인 평가를 지원합니다. 월드 반응성 수준에서는 입력으로 명시적으로 지정되지 않은 장면 조건에 따른 반응과 목표 지향적 행동을 평가합니다. 20개의 대표 모델을 분석한 결과, 성능 차이가 명확하게 나타났습니다. 카메라 기반 모델은 카메라 제어 능력에서 뛰어난 반면, 상호 작용 기능을 지원하지 않습니다. 액션 기반 모델은 객체 제어 능력이 뛰어나지만, 종종 세계의 반응성이 부족합니다. 언어 기반 모델은 상호 작용 측면에서는 더 나은 성능을 보이지만, 복잡한 제어를 덜 정확하게 따르는 경향이 있습니다. 어떤 모델도 광범위한 작업 적용 범위와 동시에 일관된 높은 성능을 보여주지 못하며, 이는 높은 시각적 품질과 명시적인 명령어 수행 능력이 월드 모델의 고유한 반응성을 보장하지 않는다는 것을 의미합니다.

Original Abstract

Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inherent reactivity of the worlds they depict: the ability to infer from the scene state how the world should react and to generate plausible consequences not explicitly described in the input. Yet existing benchmarks mainly assess visual quality or explicit instruction fulfillment by checking whether requested actions and interaction outcomes are realized, leaving inherent reactivity underexamined. We introduce WorldExam, a hierarchical diagnostic benchmark spanning four levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity. It comprises 1,474 cases across eight dedicated tasks and supports unified evaluation of camera-, action-, and language-driven model paradigms. The World Reactivity level evaluates scene-conditioned reactions and goal-directed behaviors beyond what is explicitly specified in the input. Evaluation of 20 representative models reveals a clear capability split. Camera-driven models excel at camera control, but their interfaces do not support dynamic interaction; action-driven models control subjects more precisely but often leave the world unresponsive; and language-driven models perform better on interaction but follow complex controls less faithfully. No model combines broad task coverage with consistently strong performance, showing that high visual quality and explicit instruction fulfillment do not guarantee inherent reactivity.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!