2606.17511v1 Jun 16, 2026 cs.RO

MagicSim: 실행 가능한 구체화된 상호 작용을 위한 통합 인프라

MagicSim: A Unified Infrastructure for Executable Embodied Interaction

Yuran Wang
Yuran Wang
Citations: 120
h-index: 5
Ruihai Wu
Ruihai Wu
Citations: 25
h-index: 3
Yue Chen
Yue Chen
Citations: 180
h-index: 8
Haoran Lu
Haoran Lu
Citations: 223
h-index: 6
Han Liu
Han Liu
Citations: 21
h-index: 3
Songlin Liu
Songlin Liu
Citations: 109
h-index: 1
Guo Ye
Guo Ye
Citations: 128
h-index: 5
Mutian Shen
Mutian Shen
Citations: 0
h-index: 0
Shuyang Yu
Shuyang Yu
Citations: 0
h-index: 0
Yushuang Xiao
Yushuang Xiao
Citations: 0
h-index: 0
Jihai Zhao
Jihai Zhao
Citations: 6
h-index: 2
Shang Wu
Shang Wu
Citations: 20
h-index: 2
Jianshu Zhang
Jianshu Zhang
Citations: 143
h-index: 4
Xian Gui
Xian Gui
Citations: 0
h-index: 0
Chuye Hong
Chuye Hong
Citations: 92
h-index: 3
Maojiang Su
Maojiang Su
Citations: 90
h-index: 6
Jiayi Wang
Jiayi Wang
Citations: 0
h-index: 0
Zhaoran Wang
Zhaoran Wang
Citations: 6
h-index: 2

로봇 학습 및 구체화된 에이전트는 이제 제어, 기술 및 계획을 연결하는 공유 실행 기반으로 사용될 뿐만 아니라 렌더링 도구, 컨트롤러 테스트 환경 또는 고정된 작업 환경으로 활용됩니다. 기존 파이프라인은 이러한 계층을 "마법" 액션을 통해 분리하거나, 독립적인 학습 환경을 사용하거나, 동일한 에피소드를 재현, 평가 및 주석 처리할 수 없는 단방향 렌더링만 제공합니다. 본 논문에서는 하나의 결정론적 배치 실행 환경과 공유된 마르코프 의사 결정 프로세스(MDP)를 기반으로 구축된 구체화된 상호 작용 인프라 MagicSim을 소개합니다. MagicSim은 YAML 기반의 사양을 통해 콘텐츠, 배치, 동작 및 에이전트 노출을 분리하여 작업 유형, 상호 작용 방식, 물리 법칙, 레이아웃, 센서, 아바타 및 로봇 구현 등 다양한 실행 가능한 환경을 하나의 재설정-단계 루프로 구성합니다. 공통 실행 인터페이스는 고수준 명령을 컨트롤러, 원자 기술, 계획 기본 요소 및 비동기적 계획을 통해 현실적인 로봇 액션으로 변환하며, 시뮬레이터 측면의 상태 변경이 아닌 실제 로봇 동작으로 구현됩니다. 하나의 작업 정의는 세 가지 기능을 지원합니다. 즉, 벤치마크 및 강화 학습 평가, 명령을 기반 경로로 자동 변환하는 인터페이스, 그리고 에이전트/VLM(Vision-Language Model)과 상호 작용할 수 있는 환경입니다. 자동 실행의 경우, 명령은 Command->Skill->Planner->Robot->Record 파이프라인을 통해 흐르며, 각 환경별 명령, 기술, 계획, 재시도, 주석 및 에피소드 상태는 공유된 물리 시간 단계를 기준으로 독립적으로 진행됩니다. 성공적인 실행 결과는 언어적 감독 정보, 액션 표현, 시각/기하학적 표현 및 작업 수준 상태를 포함하는 구조화된 다중 모달 경로로 저장되어, 실제 실행된 에피소드와 일치합니다. 따라서 MagicSim은 다양한 환경 구축, 구체화된 실행, 작업 평가, 자동 실행 생성 및 대화형 에이전트 인터페이스를 하나의 계획 기반 실행 환경에서 통합합니다.

Original Abstract

Robot learning and embodied agents now require simulation to serve as a shared execution substrate linking control, skills, and planning, not only as a renderer, controller testbed, or fixed task environment. Existing pipelines split these layers with "magic" actions, disconnected training environments, or forward-only renders that cannot reproduce, evaluate, and annotate the same episode. We present MagicSim, an embodied interaction infrastructure built around one deterministic batched runtime and a shared Markov decision process (MDP). From YAML-first specifications that decouple contents, placement, behavior, and agent exposure, MagicSim constructs diverse executable worlds spanning task families, interaction regimes, physics, layouts, sensors, avatars, and robot embodiments in one reset-and-step loop. A common execution interface grounds high-level commands through controllers, atomicskills, planner primitives, and asynchronous planning, realizing them as robot actions rather than simulator-side state edits. One task definition supports three capabilities: benchmark and RL evaluation, an autocollect interface that automatically turns commands into grounded trajectories, and agent/VLM-facing interaction. For automatic execution, commands flow through a Command->Skill->Planner->Robot->Record pipeline, while per-environment command, skill, planning, retry, annotation, and episode states advance independently above the shared physics tick. Successful rollouts are saved as structured multimodal trajectories aligning language supervision, action representations, visual/geometric representations, and task-level status with the executed episode. MagicSim thus unifies diverse world construction, embodied execution, task evaluation, automatic rollout generation, and interactive agent interfaces in one planner-in-the-loop runtime.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!