2606.31672v1 Jun 30, 2026 cs.CV

WorldRoamBench: 상호작용 환경 모델의 장기 안정성을 위한 개방형 벤치마크

WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models

Zhaoxu Sun
Zhaoxu Sun
Citations: 32
h-index: 3
Yang Gao
Yang Gao
Citations: 148
h-index: 5
Tingbing Xu
Tingbing Xu
Citations: 0
h-index: 0
Jiacheng Sui
Jiacheng Sui
Citations: 8
h-index: 1
Zhe Gao
Zhe Gao
Citations: 19
h-index: 2
Kewei Shi
Kewei Shi
Citations: 0
h-index: 0
Zhicheng Liu
Zhicheng Liu
Citations: 1
h-index: 1
Mingchao Sun
Mingchao Sun
Citations: 13
h-index: 2
Hongyu Pan
Hongyu Pan
Citations: 155
h-index: 5
Fan Jiang
Fan Jiang
Citations: 134
h-index: 4
Mu Xu
Mu Xu
Citations: 136
h-index: 4
Qi Fan
Qi Fan
Citations: 24
h-index: 3
Yong Li
Yong Li
Citations: 0
h-index: 0
Baoquan Chen
Baoquan Chen
Citations: 202
h-index: 5
Wenjing Yang
Wenjing Yang
Citations: 0
h-index: 0

상호작용 환경 모델(IWM) 분야에서 빠른 발전이 이루어지고 있음에도 불구하고, 기존 벤치마크는 주로 행동 추적을 경로 수준에서 평가하고 메모리와 상호 작용의 물리적 요소를 간과합니다. 본 연구에서는 IWM의 장기 안정성을 평가하기 위한 개방형 벤치마크인 WorldRoamBench를 소개합니다. 이 벤치마크는 네 가지 차원을 포함하며, 각 차원마다 다음과 같은 혁신적인 방법을 사용합니다: (i) 행동(Action): 프레임 단위 행동 지표를 사용하여 모델 간의 의미론적 규모 차이로 인해 숨겨지는 오류를 드러냅니다; (ii) 시각(Vision): 세그먼트 기반 드리프트 지표를 사용하여 시작과 끝 부분 비교만으로는 감지할 수 없는 중간 단계에서의 문제를 파악합니다; (iii) 물리(Physics): 역학, 광학 및 3차원 일관성에 대한 제어 가능성 기반 평가를 통해 실제 행동에 따른 타당성을 측정합니다; (iv) 메모리(Memory): 행동과 독립적인 프로토콜을 사용하여 장면 메모리는 전환 위치와 관련된 3차원 포인트 클라우드 재구성을 통해, 대상 메모리는 추적 및 VLM(Vision-Language Model) 추론을 통해 평가합니다. 본 벤치마크는 자연, 도시, 실내 환경의 600개 이상의 테스트 케이스를 포함하며, 1인칭 또는 3인칭 시점에서 WASD 키를 사용하여 10~60초 동안 지속적인 상호 작용이 가능합니다. 10개 이상의 공개/비공개 모델을 평가한 결과, 어떤 모델도 모든 차원에서 안정적으로 작동하지 않았으며, 가장 좋은 성능의 모델조차도 중간 정도의 점수에 그쳤습니다. WorldRoamBench를 활용한 발전은 안정적이고, 물리적인 제약 조건을 따르며, 메모리 유지 능력이 뛰어나고 실제 응용 분야에 적용 가능한 IWM 개발을 위한 중요한 단계입니다.

Original Abstract

Despite rapid progress in interactive world models (IWMs), existing benchmarks evaluate action following only at trajectory level and ignore memory and interaction physics. We introduce WorldRoamBench, an open-world benchmark for long-horizon stability across four dimensions, each with tailored innovations: (i) Action: per-frame action metric bypassing cross-model semantic scale disparity and exposing failures hidden by trajectory; (ii) Vision: segment-based drift metric capturing non-monotonic mid-sequence collapse missed by start-vs-end comparisons; (iii) Physics: controllability-gated evaluation over mechanics, optics, and 3D consistency, scoring plausibility under faithful action execution; (iv) Memory: action-decoupled protocol evaluating scene memory via transition-localized 3D point-cloud reconstruction and subject memory via tracking-plus-VLM reasoning. The benchmark comprises 600+ test cases across Nature, Urban, and Indoor scenes in first/third-person views with WASD 10-60s continuous interaction. Evaluating 10+ open/closed-source models reveals none reliably satisfies all dimensions; even the best achieves only moderate scores. Advances on WorldRoamBench are steps toward IWMs that are stable, physically grounded, memory-faithful, and deployable in real-world applications.

1 Citations
0 Influential
2.5 Altmetric
13.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!