RoboReact: 생성된 일인칭 동영상을 활용한 에이전트 기반 기술 증류를 통한 일반화 가능한 전신 조작
RoboReact: Agentic Skill Distillation from Generated Egocentric Videos for Generalizable Whole-Body Manipulation
휴머노이드 로봇은 인간 환경에서 정교한 조작을 수행할 잠재력을 가지고 있지만, 다양한 기술을 습득하는 것은 고가의 하드웨어 데이터 수집과 노동 집약적인 어노테이션 때문에 비용이 많이 든다는 문제가 있습니다. 최근 비디오 생성 모델의 발전은 시각적 관찰로부터 풍부한 조작 경험을 합성할 수 있는 유망한 기회를 제공하지만, 이러한 가상 행동을 실행 가능한 전신 휴머노이드 기술로 변환하는 것은 아직 충분히 연구되지 않았습니다. 본 논문에서는 단일 일인칭 RGB-D 데이터를 기반으로 자동으로 전신 휴머노이드 조작 기술을 합성하는 프레임워크인 RoboReact를 제시합니다. RoboReact는 인간의 조작 동영상을 생성하고, 깊이 정보를 활용한 3차원 재구성을 통해 기하학적 정보가 보존된 주요 프레임을 추출하며, 이를 고 자유도 휴머노이드 플랫폼에 적용하여 손-물체 상호 작용의 기하학적 구조를 유지합니다. RoboReact는 가상 계획과 실제 실행 간의 격차를 해소하기 위해 온라인 객체 중심 재정렬을 수행하고, 시각-언어 모델 기반의 개선 루프를 활용하여 기하학적 불일치 및 실행 오류에 대한 기술 적응 기능을 제공합니다. 개선된 기술은 전신 제어기를 통해 실행되며, 이를 통해 조화로운 전신 조작과 정교한 상호 작용이 가능합니다. 실제 휴머노이드 로봇에서의 실험 결과는 RoboReact가 다양한 물체 구성에서 일반화 성능을 보이며, 원격 제어 또는 인간의 시연 없이도 실행 중 발생하는 오류로부터 안정적으로 회복될 수 있음을 보여줍니다. 이러한 결과는 생성 모델, 시각-언어 추론 및 폐루프 제어를 결합하여 휴머노이드 기술 확보를 확장할 수 있는 잠재력을 강조합니다.
Humanoid robots have the potential to perform dexterous manipulation in human environments, yet acquiring diverse and generalizable skills remains costly due to expensive hardware data collection and labor-intensive annotation. Recent advances in video generative models provide a promising opportunity to synthesize rich manipulation experiences from visual observations, but transferring such imagined behaviors into executable whole-body humanoid skills remains largely unexplored. In this work, we present RoboReact, a framework that automatically synthesizes whole-body humanoid manipulation skills from a single egocentric RGB-D observation. RoboReact generates human manipulation videos, extracts geometry-preserving interaction keyframes through depth-aware 3D reconstruction, and retargets them to high-DoF humanoid platforms while preserving hand-object interaction geometry. To bridge the gap between imagined plans and physical execution, RoboReact performs online object-centric re-grounding and leverages a vision-language model-guided refinement loop to adapt skills under geometric mismatch and execution deviations. The refined skills are executed through a whole-body controller, enabling coordinated whole-body manipulation and dexterous interaction. Experiments on real humanoid robots demonstrate that RoboReact generalizes across diverse object configurations and robustly recovers from execution disturbances without requiring teleoperation or human demonstrations. These results highlight the potential of combining generative models, vision-language reasoning, and closed-loop control for scalable humanoid skill acquisition.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.