2607.19190v1 Jul 21, 2026 cs.RO

에이전트 기반의 Real2Sim: 비전-언어 에이전트를 활용한 물리 기반 환경 모델링

Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents

Chenfanfu Jiang
Chenfanfu Jiang
Citations: 1,369
h-index: 16
Alan L. Yuille
Alan L. Yuille
Citations: 1,063
h-index: 3
Yunuo Chen
Yunuo Chen
Citations: 43
h-index: 4
Chao Liu
Chao Liu
Citations: 0
h-index: 0
Guanxiong Chen
Guanxiong Chen
Citations: 13
h-index: 2
Qianjun Xia
Qianjun Xia
Citations: 0
h-index: 0
Jiawei Peng
Jiawei Peng
Citations: 23
h-index: 3
Heng Zhang
Heng Zhang
Citations: 0
h-index: 0
Bole Ma
Bole Ma
Citations: 0
h-index: 0
Justin Qian
Justin Qian
Citations: 0
h-index: 0
Ziyi Jiao
Ziyi Jiao
Citations: 0
h-index: 0
Bingyang Zhou
Bingyang Zhou
Citations: 43
h-index: 2
Luoxin Ye
Luoxin Ye
Citations: 135
h-index: 5
Kaifeng Zhang
Kaifeng Zhang
Citations: 234
h-index: 5
Kunyi Wang
Kunyi Wang
Citations: 29
h-index: 2
Weijia Zeng
Weijia Zeng
Citations: 199
h-index: 3
Pengzhi Yang
Pengzhi Yang
Citations: 75
h-index: 5
Ziqiu Zeng
Ziqiu Zeng
Citations: 6
h-index: 1
Huamin Wang
Huamin Wang
Citations: 33
h-index: 2
Fan Shi
Fan Shi
Citations: 10
h-index: 2
Changxi Zheng
Changxi Zheng
Citations: 19
h-index: 1
Yunzhu Li
Yunzhu Li
Citations: 229
h-index: 4
P. Y. Chen
P. Y. Chen
Citations: 86
h-index: 4
Siyuan Luo
Siyuan Luo
Citations: 24
h-index: 3

로봇과 객체 간 상호작용을 위한 실세계 데이터를 시뮬레이션 환경으로 변환하는 과정은 여전히 많은 노력이 필요합니다. 이는 단순한 시각적 재구성을 넘어, 장면의 기하 구조와 객체의 상태를 복원하고, 물리적인 파라미터를 추론하며, 액터, 객체, 카메라, 자세 및 경로를 조립하여 실행 가능한 물리 시뮬레이션을 구성해야 하기 때문입니다. 현재 이 과정은 여전히 시각 기반 모델의 수동 조정, 메시 데이터 정리, 좌표계 정렬, 그리고 다양한 시각 인식 도구와 시뮬레이터 간의 불안정한 작업 흐름에 의존하고 있습니다. 본 논문에서는 비전-언어 에이전트를 활용하여 일반화된 물리 환경 모델링을 위한 프레임워크인 extit{Agentic Real2Sim}을 소개합니다. 이 프레임워크는 객체-로봇 상호작용의 실세계 기록을 시뮬레이션 가능한 형식으로 변환하며, 관측 데이터, 기하 구조, 로봇과의 상호작용 및 객체의 상태를 보존합니다. Agentic Real2Sim은 일반적으로 별도의 Real2Sim 파이프라인으로 처리되는 강체 조작, 변형 가능한 객체와의 상호작용, 그리고 인간형 동작 장면 등 다양한 영역에서 평가되었습니다. 이는 확장 가능한 변환을 향한 첫걸음입니다. 프레임워크의 에이전트 기반 의사 결정은 최첨단 모델보다 훨씬 저렴한 비용으로 운영되는 공개 가중치 VLM 백엔드에 의해 구동되며, 동시에 유사한 수준의 변환 성공률을 달성합니다. 우리는 생성된 실세계와 일관성을 갖는 시뮬레이션 데이터를 활용하여 정책 학습 및 평가를 포함한 후속 로봇 관련 작업을 수행하고자 합니다. 프로젝트 웹사이트는 https://ericchen321.github.io/agentic_real2sim.github.io/ 에서 확인할 수 있습니다.

Original Abstract

Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a streamlined real2sim process must recover scene geometries and object states, infer physical parameters, and assemble actors, objects, cameras, poses, and trajectories into a runnable physical simulation. Today this process still depends on manual tuning of visual foundation models, mesh cleanup, coordinate-frame alignment, and brittle workflow glue across visual perception tools and simulators. We introduce \textit{Agentic Real2Sim}, a framework for generalized physical world modeling with vision-language agents, converting a real-world recording of object-robot interaction into a simulatable episodic twin which preserves observations, geometries, robot interactions, and object states. We evaluate Agentic Real2Sim on rigid-object manipulation, deformable-object interaction, and humanoid motion scenes, spanning domains that are usually handled by separate Real2Sim pipelines, marking a first step toward scalable conversion. The framework's agentic decisions can be driven by an open-weight VLM backend at a small fraction of the cost of frontier models, while attaining comparable conversion success rate. We aim to use the resulting real-world-aligned twins for downstream robotics tasks, specifically policy learning and evaluation. The project site is available at https://ericchen321.github.io/agentic_real2sim.github.io/.

0 Citations
0 Influential
8 Altmetric
40.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!