2608.06257v1 Aug 06, 2026 cs.CV

MASS: 권위 있는 공유 상태를 갖는 멀티플레이어 월드 모델

MASS: Multiplayer World Models with Authoritative Shared State

Kaipeng Zhang
Kaipeng Zhang
Citations: 137
h-index: 5
Ziqi Cai
Ziqi Cai
Citations: 33
h-index: 2
Siqi Yang
Siqi Yang
Citations: 41
h-index: 4
Yimu Wang
Yimu Wang
Citations: 0
h-index: 0
Zixian Gao
Zixian Gao
Citations: 0
h-index: 0
Yunheng Liu
Yunheng Liu
Citations: 16
h-index: 1
Shuchen Weng
Shuchen Weng
Citations: 573
h-index: 13
Erwin Wu
Erwin Wu
Citations: 9
h-index: 2
Boxin Shi
Boxin Shi
Citations: 319
h-index: 10

현재 비디오 기반 월드 모델은 멀티플레이어 환경에서 어려움을 겪는데, 이는 월드 상태와 시점에 따라 달라지는 시각적 정보가 결합되어 계산량 증가, 시점 불일치 및 확장성 저하를 야기하기 때문입니다. 본 논문에서는 이러한 제한점을 해결하기 위해 권위 있는 공유 상태(Authoritative Shared State)를 갖는 멀티플레이어 월드 모델인 MAS(Multiplayer world models with Authoritative Shared State)를 제안합니다. MAS는 멀티플레이어 게임 아키텍처에서 영감을 받아 월드 동역학과 렌더링을 분리합니다. 학습된 로직 엔진은 수동으로 작성된 전환 함수 없이, 여러 에이전트의 공동 행동으로부터 전역적이고 권위 있는 타입화된 상태를 업데이트하며, 이는 유일한 순환 메모리와 동기화 기준 역할을 합니다. 이 공유 상태로부터 학습된 렌더링 엔진은 요청 시 어떤 카메라에서도 독립적이고 일관성 있는 뷰를 생성합니다. 이러한 명시적인 분리는 MAS가 최첨단 멀티뷰 모델보다 우수한 상태 정확도를 달성하고 교차 뷰 불일치를 줄이는 데 기여합니다. 또한, MAS는 1,024명의 동시 플레이어를 대상으로 1만 번의 반복 단계를 예측하는 데 사용될 수 있습니다. 실험 결과는 명시적이고 권위 있는 상태 모델링이 확장 가능하고 일관성 있는 멀티 에이전트 월드 시뮬레이션을 위한 실용적인 기반을 제공한다는 것을 보여줍니다.

Original Abstract

Current video world models struggle in multiplayer environments because they entangle world state with view-dependent visual latents, leading to redundant compute, view inconsistencies, and poor scalability. We propose MASS (Multiplayer world models with Authoritative Shared State) to resolve this limitation. Inspired by multiplayer game architectures, MASS disentangles world dynamics and view rendering. A learned Logic Engine advances a global, authoritative typed state from joint actions without any hand-written transition function, acting as the sole recurrent memory and synchronization reference. From this shared state, a learned Rendering Engine generates independent and consistent views for any requested camera on demand. This explicit disentangling allows MASS to achieve superior state accuracy and lower cross-view inconsistency compared to state-of-the-art multi-view baselines on a matched multiplayer Snake benchmark. It advances predicted worlds with 1,024 concurrent players for 10,000 recurrent steps. Our results show that explicit, authoritative state modeling provides a practical foundation for scalable and consistent multi-agent world simulation.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!