ABot-World-0: 단일 데스크톱 GPU를 활용한 무한 인터랙티브 월드 생성 시스템
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
본 논문에서는 실시간, 장기적인 폐루프 상호작용을 지원하는 액션 기반 비디오 월드 모델인 ABot-World-0을 제시합니다. 이 모델은 AAA 게임, 시뮬레이션 엔진 및 인터넷 비디오를 포함한 다양한 데이터 인프라를 활용하여 제어 가능한 월드 역학을 학습합니다. WorldExplorer는 훈련 피드백에 의해 안내되는 에이전트 기반 수집 기능을 수행하며, 통일된 파이프라인은 14가지 결정적 품질 검사, VLM(Vision-Language Model) 기반 평가 및 동기화된 액션 및 텍스트 주석을 적용합니다. 우리는 teacher forcing과 ODE(Ordinary Differential Equation) 증류를 통해 양방향 액션 기반 가이드 모델을 원인 관계에 기반한 학생 모델로 점진적으로 증류하며, LongForcing 기법을 도입하여 장기적인 학생 모델의 자체 생성을 확장된 시야를 가진 가이드 모델과 일치시켜 누적된 분포 변화 및 자기회귀 드리프트를 완화합니다. 키보드 액션을 활용하여 장면 탐색 및 3인칭 캐릭터 상호 작용을 위한 통일된 제어 인터페이스를 제공하며, 참조 캐릭터 메모리는 3인칭 생성 과정에서 지속적인 시각적 단서를 제공하여 동일성 일관성을 유지합니다. 배포를 위해 경량 VAE(Variational Autoencoder) 디코더, 효율적인 어텐션, 메모리 기반 스케줄링 및 저비트 DiT(Diffusion Transformer) 추론을 결합한 스트리밍 추론 시스템을 공동 설계했습니다. 최적화된 저비트 구성에서 ABot-World-0은 단일 NVIDIA RTX 5090 데스크톱 GPU에서 최대 16 FPS로 720P 비디오를 스트리밍하며, 액션부터 첫 프레임까지의 지연 시간은 약 1.2초이고 피크 VRAM 사용량은 약 19GiB입니다. WorldRoamBench 및 확장된 인터랙티브 생성 실험 결과는 ABot-World-0이 우수한 제어 가능성과 일관성 있는 장기적인 월드 진화를 제공함을 보여줍니다.
We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-source data infrastructure spanning AAA games, simulation engines, and internet videos to learn controllable world dynamics. WorldExplorer performs agent-driven collection guided by training feedback, while a unified pipeline applies 14 deterministic quality checks, VLM-based assessment, and synchronized action and text annotation. We progressively distill a bidirectional action-conditioned teacher into a causal student through teacher forcing and ODE distillation, and introduce LongForcing to align long student self-rollouts with an extended-horizon teacher, mitigating accumulated distribution shift and autoregressive drift. Raw keyboard actions provide a unified control interface for scene roaming and third-person character interaction, while reference-character memory provides persistent appearance cues for identity consistency during third-person rollouts. For deployment, we co-design a streaming inference stack with a lightweight VAE decoder, efficient attention, memory-aware scheduling, and low-bit DiT inference. Across optimized low-bit configurations, ABot-World-0 streams 720P video at up to 16 FPS on a single NVIDIA RTX 5090 desktop GPU, with 1.2s action-to-first-frame latency and approximately 19GiB peak VRAM. Experiments on WorldRoamBench and extended interactive rollouts demonstrate competitive controllability and coherent long-horizon world evolution.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.