SAFECAST: 콘트라스트 집합 기반 학습 및 교정을 통한 VLA 정책의 강력한 오류 감지
SAFECAST: Robust Failure Detection for VLA Policies with Contrast-Set Training and Calibration
시각-언어-행동(VLA) 정책은 종종 배포 시 발생하는 환경 변화, 예를 들어 혼잡, 방해 객체, 조명 변화, 새로운 객체, 변경된 초기 상태 및 수정된 지침으로 인해 실패합니다. 은닉 상태 기반 위험 탐지 기술과 함수적 컨포멀 예측을 결합하여 롤아웃 오류를 감지할 수 있지만, 이러한 기술의 신뢰성은 교정 데이터가 배포 환경 조건을 정확하게 반영하는지에 따라 달라집니다. 본 연구에서는 SAFECAST를 제안하며, 이는 콘트라스트 집합 기반 변환을 활용하여 은닉 상태 탐지 기술의 학습 및 교정을 개선하고, 배포 시 발생하는 환경 변화에 대한 강건성을 향상시킵니다. 실험 결과, SAFECAST는 실제 DROID 및 LIBERO 시뮬레이션 환경에서 다양한 VLM(Vision-Language Model) 구조를 사용하여 수행한 여러 실험에서 최첨단 기준 모델보다 통계적으로 유의미하게 더 높은 오류 감지 ROC-AUC 점수를 달성했습니다. 또한, 시각적 및 언어적 콘트라스트 집합 기반 변환을 모두 사용할 때 데이터 증강 효과가 가장 크며, 콘트라스트 집합 기반 변환을 통해 얻은 데이터를 사용하여 시뮬레이션에서 실제 환경으로의 교정(sim-to-real calibration)이 실제 롤아웃 데이터를 사용하는 것보다 더 나은 성능을 보이는 것을 확인했습니다.
Vision-language-action policies often fail under deployment-time distribution shifts such as clutter, distractor objects, lighting changes, novel objects, altered initial states, and reworded instructions. Hidden-state-based risk probes combined with functional conformal prediction can detect rollout failures, but their reliability depends on calibration data matching deployment conditions. We introduce SAFECAST, which leverages contrast set perturbations to improve hidden-state probe training and calibration for deployment time shift. SAFECAST statistically significantly improves failure detection ROC-AUC scores over a state of the art baseline in both real-world DROID and LIBERO simulation experiments across multiple VLM backbones. We further find that SAFECAST benefits most when both visual and language contrast set perturbations are used to augment data, and that with contrast set perturbations, sim-to-real calibration leads to better probes than using real rollout data only.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.