WorldRoamBench: 상호작용 환경 모델의 장기 안정성을 위한 개방형 벤치마크
WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models
상호작용 환경 모델(IWM) 분야에서 빠른 발전이 이루어지고 있음에도 불구하고, 기존 벤치마크는 주로 행동 추적을 경로 수준에서 평가하고 메모리와 상호 작용의 물리적 요소를 간과합니다. 본 연구에서는 IWM의 장기 안정성을 평가하기 위한 개방형 벤치마크인 WorldRoamBench를 소개합니다. 이 벤치마크는 네 가지 차원을 포함하며, 각 차원마다 다음과 같은 혁신적인 방법을 사용합니다: (i) 행동(Action): 프레임 단위 행동 지표를 사용하여 모델 간의 의미론적 규모 차이로 인해 숨겨지는 오류를 드러냅니다; (ii) 시각(Vision): 세그먼트 기반 드리프트 지표를 사용하여 시작과 끝 부분 비교만으로는 감지할 수 없는 중간 단계에서의 문제를 파악합니다; (iii) 물리(Physics): 역학, 광학 및 3차원 일관성에 대한 제어 가능성 기반 평가를 통해 실제 행동에 따른 타당성을 측정합니다; (iv) 메모리(Memory): 행동과 독립적인 프로토콜을 사용하여 장면 메모리는 전환 위치와 관련된 3차원 포인트 클라우드 재구성을 통해, 대상 메모리는 추적 및 VLM(Vision-Language Model) 추론을 통해 평가합니다. 본 벤치마크는 자연, 도시, 실내 환경의 600개 이상의 테스트 케이스를 포함하며, 1인칭 또는 3인칭 시점에서 WASD 키를 사용하여 10~60초 동안 지속적인 상호 작용이 가능합니다. 10개 이상의 공개/비공개 모델을 평가한 결과, 어떤 모델도 모든 차원에서 안정적으로 작동하지 않았으며, 가장 좋은 성능의 모델조차도 중간 정도의 점수에 그쳤습니다. WorldRoamBench를 활용한 발전은 안정적이고, 물리적인 제약 조건을 따르며, 메모리 유지 능력이 뛰어나고 실제 응용 분야에 적용 가능한 IWM 개발을 위한 중요한 단계입니다.
Despite rapid progress in interactive world models (IWMs), existing benchmarks evaluate action following only at trajectory level and ignore memory and interaction physics. We introduce WorldRoamBench, an open-world benchmark for long-horizon stability across four dimensions, each with tailored innovations: (i) Action: per-frame action metric bypassing cross-model semantic scale disparity and exposing failures hidden by trajectory; (ii) Vision: segment-based drift metric capturing non-monotonic mid-sequence collapse missed by start-vs-end comparisons; (iii) Physics: controllability-gated evaluation over mechanics, optics, and 3D consistency, scoring plausibility under faithful action execution; (iv) Memory: action-decoupled protocol evaluating scene memory via transition-localized 3D point-cloud reconstruction and subject memory via tracking-plus-VLM reasoning. The benchmark comprises 600+ test cases across Nature, Urban, and Indoor scenes in first/third-person views with WASD 10-60s continuous interaction. Evaluating 10+ open/closed-source models reveals none reliably satisfies all dimensions; even the best achieves only moderate scores. Advances on WorldRoamBench are steps toward IWMs that are stable, physically grounded, memory-faithful, and deployable in real-world applications.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.