표현 오토인코더를 활용한 멀티플레이어 인터랙티브 월드 모델
Multiplayer Interactive World Models with Representation Autoencoders
본 논문에서는 복잡한 물리적 상호작용으로 인해 매우 역동적인 환경에 대한 최초의 멀티플레이어 월드 모델을 소개합니다. 기존의 싱글 플레이어 월드 모델이 다른 에이전트를 환경의 일부로 취급하는 반면, 저희의 모델은 여러 에이전트의 행동 스트림에 기반하여 학습하며, 장면의 변화를 올바른 플레이어에게 귀속시키고, 플레이어들의 임의적인 행동 조합 하에서도 일관성을 유지합니다. 본 연구는 빠른 속도와 밀접하게 결합된 역학 관계를 가진 게임인 락앤롤 레전드(Rocket League)에서 이 문제를 다룹니다. 공개적으로 사용 가능한 봇을 사용하여 수집된 10,000시간의 게임 플레이 데이터를 기반으로 학습된 저희의 50억 파라미터 규모의 잠재 확산 모델은 실시간으로 4인용 매치를 생성하며, 단일 Nvidia B200 GPU에서 초당 20프레임을 출력합니다. 짧은 클립 데이터로만 학습되었음에도 불구하고, 이 모델의 예측 결과는 학습 범위를 훨씬 넘어 안정적으로 유지됩니다. 측정 가능한 가장 긴 기간 동안인 5분까지 분포 품질이 일정하게 유지되며, 실제로는 모델이 몇 시간 동안 작동하면서도 붕괴되는 징후를 보이지 않습니다. 저희는 비디오 코덱, 생성 목표 및 멀티플레이어 조건 설정과 같은 핵심 설계 선택 사항을 체계적으로 조사합니다. 또한, 모델과 데이터의 규모 변화에 따른 행동 변화를 분석하고, 나타나는 능력과 지속되는 실패 요인을 파악했습니다. 더 나아가, 모델의 시각적 외관이 아닌 물리적인 이해도를 평가하기 위한 표적 평가 방법을 개발했습니다. 멀티플레이어 월드 모델에 대한 추가 연구를 지원하기 위해, 저희는 데이터셋, 전체 학습 및 추론 코드를 공개하며, 실시간 데모도 제공합니다.
We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions. Whereas single-player world models treat the other agents as part of the environment, ours conditions on the action streams of multiple agents, learning to attribute changes in the scene to the correct player and to stay coherent under arbitrary combinations of their actions. We study this problem in the game of Rocket League, where players compete and cooperate under fast, tightly coupled dynamics. Trained on 10,000 hours of gameplay collected with publicly available bots, our 5-billion-parameter latent diffusion model generates four-player matches in real time, producing 20 frames per second on a single Nvidia B200 GPU. Although trained only on short clips, its rollouts stay stable far beyond the training horizon: distributional quality holds steady out to five minutes, the longest horizon we measure, and in practice we observe rollouts continuing for hours with no sign of collapse. We systematically investigate the central design choices: the video codec, the generative objective, and the multiplayer conditioning scheme. In addition, we characterize how behavior changes with model and data scale, including the capabilities that emerge and the failure modes that persist. We further develop targeted evaluations that probe the model's physical understanding rather than visual appearance alone. To support continued research on multiplayer world models, we release our dataset, our full training and inference codebase, and a live demo.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.