CG-월드: 대규모 월드 상태 데이터셋 및 월드 모델을 위한 프로토콜
CG-World: A Large-Scale World-State Dataset and Protocol for World Models
월드 모델은 상태, 행동, 이벤트 및 관찰의 복합적인 동역학을 학습해야 하지만, 기존의 비디오, 로봇 공학 및 시뮬레이션 데이터셋은 일반적으로 이러한 구조의 일부만 포착합니다. 본 논문에서는 산업용 컴퓨터 그래픽 제작 파이프라인에서 파생된 대규모 월드 상태 데이터셋인 CG-월드를 소개합니다. CG-월드는 다중 모달 의미, 공간 구조, 골격 및 컨트롤러 상태, 운동 궤적, 카메라 및 조명 매개변수, 물리 엔진 캐시, 접촉 이벤트, 그리고 멀티 패스 렌더링을 포함한 중간 상태를 명시적으로 기록합니다. CG-월드 v1은 약 85만 개의 시간적으로 정렬된 1~5초 길이의 세그먼트를 포함하며, 잠재 상태, 관찰 결과, 관계, 이벤트 및 분기 메타데이터를 분리하여 통합된 시공간 샘플로 구성합니다. 개입 학습 및 반사실적 추론을 지원하기 위해, CG-월드는 사실적인 궤적, 관찰 개입, 행동 개입, 메커니즘 개입, 그리고 엄격한 반사실적 분기를 포함하는 분기 계보를 정의하며, 개입 대상, 불변성, 그리고 대안 결과들을 명시적으로 기록합니다. 본 데이터셋은 기하학적 조건에 따른 비디오 생성, 행동 예측 및 폐루프 시각-언어-행동 정책 전송에 대한 평가를 수행했습니다. 그 결과, CG-월드는 제어된 생성, 행동 모델링 및 임베디드 정책 전송을 위한 재사용 가능한 구조화된 감독 신호를 제공합니다. 우리는 지속적인 데이터 수집과 커뮤니티 협력을 통해 월드 모델, 물리 AI 및 임베디드 지능을 위한 공유 데이터 인프라를 구축하는 방향으로 CG-월드를 확장할 계획입니다.
World models must learn the joint dynamics of states, actions, events, and observations, yet existing video, robotics, and simulation datasets usually capture only part of this structure. We introduce CG-World, a large-scale world-state dataset and protocol derived from industrial computer graphics production pipelines. CG-World explicitly records intermediate states, including multimodal semantics, spatial structure, skeletal and controller states, motion curves, camera and lighting parameters, physics caches, contact events, and multi-pass renderings. CG-World v1 contains approximately 850,000 temporally aligned segments of 1-5 seconds. It separates latent states, observations, relations, events, and branch metadata, and organizes them into unified spatiotemporal samples. To support intervention learning and counterfactual reasoning, CG-World defines a branch lineage covering factual trajectories, observation interventions, action interventions, mechanism interventions, and strict counterfactual branches, with intervention targets, invariants, and alternative outcomes explicitly recorded. We evaluate the dataset on geometry-conditioned video generation, action prediction, and closed-loop vision-language-action policy transfer. Results show that CG-World provides reusable structured supervision for controlled generation, action modeling, and embodied policy transfer. We plan to expand CG-World through continued data collection and community collaboration toward a shared data infrastructure for world models, Physical AI, and embodied intelligence.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.