PlayWorld: 장기 목표를 가진 에이전트 플레이어를 활용한 월드 모델 성능 평가
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives
비디오 월드 모델은 현재 관찰 및 사용자 행동에 기반하여 미래 상태를 시뮬레이션합니다. 최근 시스템들은 긴 시퀀스에서 뛰어난 비디오 일관성 및 액션 제어 능력을 보여주었습니다. 그러나 이러한 인터랙티브 모델들을 공정하게 비교하는 것은 여전히 어렵습니다. 일반적으로 사용자는 월드 모델을 평가할 때, 상호작용을 통해 장기 목표를 달성합니다. 예를 들어, 사용자는 환경의 일관성을 확인하기 위해 360도 회전을 하거나, 물에 들어가 현실적인 물결이 생성되는지 확인할 수 있습니다. 동일한 목표를 달성하는 데 필요한 액션 시퀀스는 모델마다 크게 다를 수 있으므로, 고정된 액션 기반 평가 방식은 모델 간 비교에 적합하지 않습니다. 이러한 문제를 해결하기 위해, 우리는 멀티모달 에이전트 플레이어를 활용하여 월드 모델과 상호작용하며 특정 장기 목표를 달성하도록 합니다. 이러한 패러다임을 바탕으로, 우리는 171개의 시나리오로 구성된 PlayWorld라는 벤치마크를 소개합니다. 성능을 종합적으로 평가하기 위해, 우리는 기하학적 일관성, 상호 작용 충실도, 가시 범위를 벗어난 환경 변화, 그리고 추론 능력의 변화와 같은 네 가지 핵심 요소를 중심으로 모델을 평가합니다. 또한, 비디오 품질 및 제어 가능성에 대한 기본적인 지표를 포함했습니다. 9개의 최첨단 월드 모델에 대한 실험 결과, 현재 모델들은 특히 공간적 일관성을 유지하고 지속적인 상태 변화를 나타내는 데 있어 장기적인 인터랙티브 목표 달성 측면에서 여전히 신뢰성이 부족하다는 것을 보여줍니다. 코드 및 데이터는 https://github.com/kxding/PlayWorld 에서 확인할 수 있습니다.
Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long sequences. However, fairly comparing these interactive models remains challenging. In practice, a human player typically evaluates a world model by pursuing long-horizon objectives through interaction. For example, a user may turn around 360 degrees to see whether the environment remains consistent, or walk into the water and inspect whether realistic water ripples are generated. The action sequence required to achieve the same objective may vary substantially between models, making fixed action-conditioned evaluation unsuitable for cross-model comparison. To address this, we employ multi-modal Agent Players to interact with world models toward specified long-horizon objectives. Building on this paradigm, we introduce PlayWorld, a benchmark providing 171 scenarios, each with a specified objective. To evaluate performance thoroughly, we assess models along four core dimensions: geometry consistency, interaction fidelity, out-of-sight evolution, and insight evolution. In addition, we incorporate basic ability metrics for video quality and controllability. Experiments across nine state-of-the-art world models reveal that current models remain unreliable on long-horizon interactive objectives, particularly in maintaining spatial consistency and persistent state evolution. Code and data are available at https://github.com/kxding/PlayWorld.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.