2606.19980v1 Jun 18, 2026 cs.AI

ENPIRE: 실제 환경에서 로봇 정책의 능동적 자기 개선

ENPIRE: Agentic Robot Policy Self-Improvement in the Real World

Guanya Shi
Guanya Shi
Citations: 331
h-index: 9
Yi Yang
Yi Yang
Citations: 354
h-index: 4
Ken Goldberg
Ken Goldberg
Citations: 125
h-index: 5
Jiahong Xie
Jiahong Xie
Citations: 0
h-index: 0
Tonghe Zhang
Tonghe Zhang
Citations: 155
h-index: 4
Haotian Lin
Haotian Lin
Citations: 52
h-index: 2
Letian
Letian
Citations: 7
h-index: 2
“Max” Fu
“Max” Fu
Citations: 0
h-index: 0
Haoru Xue
Haoru Xue
Citations: 304
h-index: 7
Jalen Lu
Jalen Lu
Citations: 1
h-index: 1
Cunxi Dai
Cunxi Dai
Citations: 70
h-index: 4
Zi Wang
Zi Wang
Citations: 372
h-index: 6
Jimmy Wu
Jimmy Wu
Citations: 164
h-index: 2
Guanzhi Wang
Guanzhi Wang
California Institute of Technology
Citations: 5,648
h-index: 16
S. Sastry
S. Sastry
Citations: 224
h-index: 9
“Jim” Fan
“Jim” Fan
Citations: 214
h-index: 1
Yuke Zhu
Yuke Zhu
Citations: 530
h-index: 7
Cmu
Cmu
Citations: 458
h-index: 4
UC Berkeley
UC Berkeley
Citations: 61
h-index: 3
Wenli Xiao
Wenli Xiao
Citations: 1,714
h-index: 15

실제 환경에서의 숙련된 로봇 조작은 인적 감독과 알고리즘 설계에 크게 의존하며, 이는 일반적인 물리적 지능을 추구하는 데 있어 중요한 병목 현상입니다. 최근 개발된 코딩 에이전트는 알고리즘 검색을 자동화하는 코드를 생성할 수 있지만, 이러한 성공 사례는 대부분 디지털 환경으로 제한됩니다. 우리는 로봇 연구를 자동화하는 데 필요한 핵심 요소가 실제 환경에서의 정책 개선을 위한 반복적인 피드백 루프라는 가설을 세웠습니다. 즉, 장면을 초기화하고, 정책을 실행하며, 결과를 검증하고, 다음 반복을 개선하는 과정을 거치는 것입니다. 이러한 격차를 해소하기 위해, 우리는 ENPIRE라는 코딩 에이전트를 위한 프레임워크를 소개합니다. ENPIRE는 실제 환경의 피드백 루틴을 구현하며, 자동 초기화 및 검증 기능을 제공하는 환경 모듈(EN), 정책 개선을 수행하는 정책 개선 모듈(PI), 여러 물리 로봇을 동시에 사용하여 정책을 평가하는 롤아웃 모듈(R), 그리고 에이전트가 로그를 분석하고, 문헌을 참고하며, 교육 인프라 및 알고리즘 코드를 개선하여 오류 패턴을 해결하는 진화 모듈(E)의 네 가지 핵심 모듈로 구성됩니다. 이 폐쇄 루프 시스템은 실제 환경에서의 조작 학습을 제어 가능한 최적화 프로세스로 변환하여 인간의 노력을 최소화하면서 다양한 교육 방식과 에이전트 변형에 대한 공정한 실험을 가능하게 합니다. ENPIRE를 통해, 선도적인 코딩 에이전트는 핀 박스 정리, 지퍼 타이라 고정, 도구 사용 등 어려운 숙련된 조작 작업에서 99%의 성공률을 달성하는 정책을 자율적으로 학습할 수 있습니다. 또한, 로봇 군에 여러 에이전트를 배포하면 이 프로세스가 더욱 가속화됩니다. 우리의 결과는 코딩 에이전트를 활용하여 실제 세계의 로봇 기술을 자율적으로 발전시키는 실용적이고 확장 가능한 방법을 제시합니다.

Original Abstract

Achieving dexterous robotic manipulation in the real world heavily relies on human supervision and algorithm engineering, which becomes a central bottleneck in the pursuit of general physical intelligence. Although emerging coding agents can generate code to automate algorithm search, their successes remain largely confined in digital environments. We conjecture that the missing abstraction to automate robotics research is a repeatable feedback loop for real-world policy improvement: reset the scene, execute a policy, verify the outcome, and refine the next iteration. To bridge this gap, we introduce ENPIRE, a harness framework for coding agents that instantiates this physical feedback routine with four core modules: an Environment module (EN) for automatic reset and verification, a Policy Improvement module (PI) that launches policy refinement, a Rollout module (R) to evaluate policies with one or multiple physical robots operating in parallel, and an Evolution module (E) in which coding agents analyze logs, consult literature, improve training infrastructure and algorithm code to address failure modes. This closed-loop system transforms real-world manipulation learning into a controllable optimization procedure, minimizing human effort while allowing fair ablations across training recipe and agent variants. Powered by ENPIRE, frontier coding agents can autonomously train a policy to achieve a 99% success rate on challenging, dexterous manipulation tasks, such as organizing a pin box, fastening a zip tie, and tool use, a process that further accelerates when we dispatch an agent team on a robot fleet. Our results suggest a practical and scalable path toward deploying coding agents to autonomously advancing robotics in the physical world.

5 Citations
1 Influential
8 Altmetric
47.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!