2604.16683v1 Apr 17, 2026 cs.RO

Rewind-IL: 모방 학습을 위한 온라인 오류 감지 및 상태 복구 기술

Rewind-IL: Online Failure Detection and State Respawning for Imitation Learning

Gehan Zheng
Gehan Zheng
Citations: 0
h-index: 0
S.N. Seenivasan
S.N. Seenivasan
Citations: 0
h-index: 0
Matthew Johnson-Roberson
Matthew Johnson-Roberson
Citations: 345
h-index: 9
Weiming Zhi
Weiming Zhi
Citations: 193
h-index: 9

모방 학습은 로봇이 시연 데이터를 통해 복잡한 시각-운동 조작 기술을 습득하도록 하는 데 기여했지만, 특히 장기적인 행동 단위를 사용하는 정책의 경우, 실제 적용 과정에서의 오류 발생은 여전히 중요한 문제점으로 남아 있습니다. 실행이 시연 데이터 범위를 벗어나면, 이러한 정책은 종종 오류를 회복하지 못하고 여전히 현지적으로 타당한 행동을 생성하는 경향이 있습니다. 기존의 런타임 모니터는 오류 데이터를 필요로 하거나, 정상적인 상황에서도 과도하게 오류를 감지하거나, 오류를 감지하면 더 이상 복구 메커니즘을 제공하지 않습니다. 본 논문에서는 생성적인 행동 단위 기반의 모방 학습 정책을 위한 훈련이 필요 없는 온라인 안전 장치 프레임워크인 Rewind-IL을 제시합니다. Rewind-IL은 Temporal Inter-chunk Discrepancy Estimate (TIDE)를 기반으로 하는 제로샷 오류 감지기와, 분할된 컨포멀 예측을 통해 보정된 상태 복구 메커니즘을 결합합니다. 오프라인에서는 비전-언어 모델이 시연 데이터 내의 복구 지점을 식별하고, 고정된 정책 인코더를 사용하여 간결한 복구 지점 특징 데이터베이스를 구축합니다. 온라인에서는 Rewind-IL이 겹치는 행동 단위의 일관성을 모니터링하고, 복구 지점 라이브러리와의 유사성을 추적하며, 오류가 발생하면 가장 최근에 확인된 안전한 상태로 실행을 되돌려 정책의 초기 상태에서 추론을 다시 시작합니다. 실제 환경 및 시뮬레이션 환경에서 장기 조작 작업에 대한 실험 결과, 정책 내부의 일관성과 의미론적으로 기반한 상태 복구가 모방 학습의 신뢰성을 향상시키는 실용적인 방법임을 보여줍니다. 추가 자료는 https://sjay05.github.io/rewind-il 에서 확인할 수 있습니다.

Original Abstract

Imitation learning has enabled robots to acquire complex visuomotor manipulation skills from demonstrations, but deployment failures remain a major obstacle, especially for long-horizon action-chunked policies. Once execution drifts off the demonstration manifold, these policies often continue producing locally plausible actions without recovering from the failure. Existing runtime monitors either require failure data, over-trigger under benign feature drift, or stop at failure detection without providing a recovery mechanism. We present Rewind-IL, a training-free online safeguard framework for generative action-chunked imitation policies. Rewind-IL combines a zero-shot failure detector based on Temporal Inter-chunk Discrepancy Estimate (TIDE), calibrated with split conformal prediction, with a state-respawning mechanism that returns the robot to a semantically verified safe intermediate state. Offline, a vision-language model identifies recovery checkpoints in demonstrations, and the frozen policy encoder is used to construct a compact checkpoint feature database. Online, Rewind-IL monitors self-consistency in overlapping action chunks, tracks similarity to the checkpoint library, and, upon failure, rewinds execution to the latest verified safe state before restarting inference from a clean policy state. Experiments on real-world and simulated long-horizon manipulation tasks, including transfer to flow-matching action-chunked policies, demonstrate that policy-internal consistency coupled with semantically grounded respawning offers a practical route to improved reliability in imitation learning. Supplemental materials are available at https://sjay05.github.io/rewind-il

4 Citations
0 Influential
4.5 Altmetric
26.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!