2608.04246v1 Aug 04, 2026 cs.RO

SAFECAST: 콘트라스트 집합 기반 학습 및 교정을 통한 VLA 정책의 강력한 오류 감지

SAFECAST: Robust Failure Detection for VLA Policies with Contrast-Set Training and Calibration

Abrar Anwar
Abrar Anwar
Citations: 385
h-index: 9
Jesse Thomason
Jesse Thomason
Citations: 83
h-index: 4
Harshitha Rajaprakash
Harshitha Rajaprakash
Citations: 8
h-index: 1
Aditeya Prajapati
Aditeya Prajapati
Citations: 0
h-index: 0
Rong Xue
Rong Xue
Citations: 31
h-index: 3

시각-언어-행동(VLA) 정책은 종종 배포 시 발생하는 환경 변화, 예를 들어 혼잡, 방해 객체, 조명 변화, 새로운 객체, 변경된 초기 상태 및 수정된 지침으로 인해 실패합니다. 은닉 상태 기반 위험 탐지 기술과 함수적 컨포멀 예측을 결합하여 롤아웃 오류를 감지할 수 있지만, 이러한 기술의 신뢰성은 교정 데이터가 배포 환경 조건을 정확하게 반영하는지에 따라 달라집니다. 본 연구에서는 SAFECAST를 제안하며, 이는 콘트라스트 집합 기반 변환을 활용하여 은닉 상태 탐지 기술의 학습 및 교정을 개선하고, 배포 시 발생하는 환경 변화에 대한 강건성을 향상시킵니다. 실험 결과, SAFECAST는 실제 DROID 및 LIBERO 시뮬레이션 환경에서 다양한 VLM(Vision-Language Model) 구조를 사용하여 수행한 여러 실험에서 최첨단 기준 모델보다 통계적으로 유의미하게 더 높은 오류 감지 ROC-AUC 점수를 달성했습니다. 또한, 시각적 및 언어적 콘트라스트 집합 기반 변환을 모두 사용할 때 데이터 증강 효과가 가장 크며, 콘트라스트 집합 기반 변환을 통해 얻은 데이터를 사용하여 시뮬레이션에서 실제 환경으로의 교정(sim-to-real calibration)이 실제 롤아웃 데이터를 사용하는 것보다 더 나은 성능을 보이는 것을 확인했습니다.

Original Abstract

Vision-language-action policies often fail under deployment-time distribution shifts such as clutter, distractor objects, lighting changes, novel objects, altered initial states, and reworded instructions. Hidden-state-based risk probes combined with functional conformal prediction can detect rollout failures, but their reliability depends on calibration data matching deployment conditions. We introduce SAFECAST, which leverages contrast set perturbations to improve hidden-state probe training and calibration for deployment time shift. SAFECAST statistically significantly improves failure detection ROC-AUC scores over a state of the art baseline in both real-world DROID and LIBERO simulation experiments across multiple VLM backbones. We further find that SAFECAST benefits most when both visual and language contrast set perturbations are used to augment data, and that with contrast set perturbations, sim-to-real calibration leads to better probes than using real rollout data only.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!