시각적 검증을 통한 추론 시간 정책 제어 및 자율적인 정책 개선
Visual Verification Enables Inference-time Steering and Autonomous Policy Improvement
실제 환경에 배치된 로봇은 경험으로부터 학습하고 시간이 지남에 따라 성능을 향상시켜야 합니다. 이를 위해서는 연습과 피드백 기반의 학습 메커니즘이 필요합니다. 본 논문에서는 추론 시간 정책 제어 및 자율적인 정책 개선을 위한 범용 로봇 정책 프레임워크인 VERITAS를 제안합니다. 사전 훈련된 범용 로봇 정책을 ``생성기(generator)``로 사용하고, 추론 시간에 액션을 평가하는 경사 기반이 아닌 ``시각적 검증기(visual verifier)``와 결합하여 이 프레임워크를 구성했습니다. 이 프레임워크는 추가적인 훈련 없이도 정책 성능을 향상시키는 추론 시간 제어를 가능하게 합니다. 실험 결과, 시각적 검증은 추가적인 데모 데이터에 대한 훈련 없이 기존의 범용 로봇 정책보다 일관되게 우수한 성능을 보였습니다. 또한, 검증된 실행 결과는 오프라인 정책 개선에 효과적인 지도 정보를 제공하며, 검증된 자체 생성 경로를 기반으로 미세 조정된 정책은 일관된 성능 향상을 달성했습니다. 주목할 점은, 검증된 실행 결과를 이용한 사후 훈련이 추가적인 인적 개입 없이 전문가 데모와 유사한 효율성을 가진다는 것입니다. 본 연구 결과는 추론 시간 검증이 배포 과정에서 로봇 정책을 개선하는 실용적이고 확장 가능한 메커니즘임을 보여줍니다.
Robots deployed in the real world should learn from their experience and improve over time. This requires a mechanism of practicing and learning from feedback. In this paper, we propose VERITAS, a generator-verifier framework for generalist robot policies for inference-time policy steering and self-improvement. We use a pre-trained generalist robot policy as a ``generator'' and pair it with a gradient-free ``visual verifier'' that evaluates actions at inference time. This framework enables inference-time steering that improves policy performance without additional training. We demonstrate that inference-time verification consistently outperforms vanilla generalists without training on additional demonstration data. Additionally, we demonstrate that the verified rollouts provide effective supervision for offline policy improvement: policies fine-tuned on verified self-generated trajectories achieve consistent performance gains. Notably, we find that post-training with verified rollouts achieves comparable efficiency to expert demonstrations, while requiring no human interventions. Our results highlight inference-time verification as a practical and scalable mechanism for improving robotic policies during deployment.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.