MagicSim: 실행 가능한 구체화된 상호 작용을 위한 통합 인프라
MagicSim: A Unified Infrastructure for Executable Embodied Interaction
로봇 학습 및 구체화된 에이전트는 이제 제어, 기술 및 계획을 연결하는 공유 실행 기반으로 사용될 뿐만 아니라 렌더링 도구, 컨트롤러 테스트 환경 또는 고정된 작업 환경으로 활용됩니다. 기존 파이프라인은 이러한 계층을 "마법" 액션을 통해 분리하거나, 독립적인 학습 환경을 사용하거나, 동일한 에피소드를 재현, 평가 및 주석 처리할 수 없는 단방향 렌더링만 제공합니다. 본 논문에서는 하나의 결정론적 배치 실행 환경과 공유된 마르코프 의사 결정 프로세스(MDP)를 기반으로 구축된 구체화된 상호 작용 인프라 MagicSim을 소개합니다. MagicSim은 YAML 기반의 사양을 통해 콘텐츠, 배치, 동작 및 에이전트 노출을 분리하여 작업 유형, 상호 작용 방식, 물리 법칙, 레이아웃, 센서, 아바타 및 로봇 구현 등 다양한 실행 가능한 환경을 하나의 재설정-단계 루프로 구성합니다. 공통 실행 인터페이스는 고수준 명령을 컨트롤러, 원자 기술, 계획 기본 요소 및 비동기적 계획을 통해 현실적인 로봇 액션으로 변환하며, 시뮬레이터 측면의 상태 변경이 아닌 실제 로봇 동작으로 구현됩니다. 하나의 작업 정의는 세 가지 기능을 지원합니다. 즉, 벤치마크 및 강화 학습 평가, 명령을 기반 경로로 자동 변환하는 인터페이스, 그리고 에이전트/VLM(Vision-Language Model)과 상호 작용할 수 있는 환경입니다. 자동 실행의 경우, 명령은 Command->Skill->Planner->Robot->Record 파이프라인을 통해 흐르며, 각 환경별 명령, 기술, 계획, 재시도, 주석 및 에피소드 상태는 공유된 물리 시간 단계를 기준으로 독립적으로 진행됩니다. 성공적인 실행 결과는 언어적 감독 정보, 액션 표현, 시각/기하학적 표현 및 작업 수준 상태를 포함하는 구조화된 다중 모달 경로로 저장되어, 실제 실행된 에피소드와 일치합니다. 따라서 MagicSim은 다양한 환경 구축, 구체화된 실행, 작업 평가, 자동 실행 생성 및 대화형 에이전트 인터페이스를 하나의 계획 기반 실행 환경에서 통합합니다.
Robot learning and embodied agents now require simulation to serve as a shared execution substrate linking control, skills, and planning, not only as a renderer, controller testbed, or fixed task environment. Existing pipelines split these layers with "magic" actions, disconnected training environments, or forward-only renders that cannot reproduce, evaluate, and annotate the same episode. We present MagicSim, an embodied interaction infrastructure built around one deterministic batched runtime and a shared Markov decision process (MDP). From YAML-first specifications that decouple contents, placement, behavior, and agent exposure, MagicSim constructs diverse executable worlds spanning task families, interaction regimes, physics, layouts, sensors, avatars, and robot embodiments in one reset-and-step loop. A common execution interface grounds high-level commands through controllers, atomicskills, planner primitives, and asynchronous planning, realizing them as robot actions rather than simulator-side state edits. One task definition supports three capabilities: benchmark and RL evaluation, an autocollect interface that automatically turns commands into grounded trajectories, and agent/VLM-facing interaction. For automatic execution, commands flow through a Command->Skill->Planner->Robot->Record pipeline, while per-environment command, skill, planning, retry, annotation, and episode states advance independently above the shared physics tick. Successful rollouts are saved as structured multimodal trajectories aligning language supervision, action representations, visual/geometric representations, and task-level status with the executed episode. MagicSim thus unifies diverse world construction, embodied execution, task evaluation, automatic rollout generation, and interactive agent interfaces in one planner-in-the-loop runtime.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.