에이전트 기반의 Real2Sim: 비전-언어 에이전트를 활용한 물리 기반 환경 모델링
Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents
로봇과 객체 간 상호작용을 위한 실세계 데이터를 시뮬레이션 환경으로 변환하는 과정은 여전히 많은 노력이 필요합니다. 이는 단순한 시각적 재구성을 넘어, 장면의 기하 구조와 객체의 상태를 복원하고, 물리적인 파라미터를 추론하며, 액터, 객체, 카메라, 자세 및 경로를 조립하여 실행 가능한 물리 시뮬레이션을 구성해야 하기 때문입니다. 현재 이 과정은 여전히 시각 기반 모델의 수동 조정, 메시 데이터 정리, 좌표계 정렬, 그리고 다양한 시각 인식 도구와 시뮬레이터 간의 불안정한 작업 흐름에 의존하고 있습니다. 본 논문에서는 비전-언어 에이전트를 활용하여 일반화된 물리 환경 모델링을 위한 프레임워크인 extit{Agentic Real2Sim}을 소개합니다. 이 프레임워크는 객체-로봇 상호작용의 실세계 기록을 시뮬레이션 가능한 형식으로 변환하며, 관측 데이터, 기하 구조, 로봇과의 상호작용 및 객체의 상태를 보존합니다. Agentic Real2Sim은 일반적으로 별도의 Real2Sim 파이프라인으로 처리되는 강체 조작, 변형 가능한 객체와의 상호작용, 그리고 인간형 동작 장면 등 다양한 영역에서 평가되었습니다. 이는 확장 가능한 변환을 향한 첫걸음입니다. 프레임워크의 에이전트 기반 의사 결정은 최첨단 모델보다 훨씬 저렴한 비용으로 운영되는 공개 가중치 VLM 백엔드에 의해 구동되며, 동시에 유사한 수준의 변환 성공률을 달성합니다. 우리는 생성된 실세계와 일관성을 갖는 시뮬레이션 데이터를 활용하여 정책 학습 및 평가를 포함한 후속 로봇 관련 작업을 수행하고자 합니다. 프로젝트 웹사이트는 https://ericchen321.github.io/agentic_real2sim.github.io/ 에서 확인할 수 있습니다.
Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a streamlined real2sim process must recover scene geometries and object states, infer physical parameters, and assemble actors, objects, cameras, poses, and trajectories into a runnable physical simulation. Today this process still depends on manual tuning of visual foundation models, mesh cleanup, coordinate-frame alignment, and brittle workflow glue across visual perception tools and simulators. We introduce \textit{Agentic Real2Sim}, a framework for generalized physical world modeling with vision-language agents, converting a real-world recording of object-robot interaction into a simulatable episodic twin which preserves observations, geometries, robot interactions, and object states. We evaluate Agentic Real2Sim on rigid-object manipulation, deformable-object interaction, and humanoid motion scenes, spanning domains that are usually handled by separate Real2Sim pipelines, marking a first step toward scalable conversion. The framework's agentic decisions can be driven by an open-weight VLM backend at a small fraction of the cost of frontier models, while attaining comparable conversion success rate. We aim to use the resulting real-world-aligned twins for downstream robotics tasks, specifically policy learning and evaluation. The project site is available at https://ericchen321.github.io/agentic_real2sim.github.io/.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.