ActFovea: 시공간적 시각-행동 일관성을 활용한 VLA 정책의 실시간 보호
ActFovea: Runtime Safeguarding for VLA Policies via Spatiotemporal Visual-Action Consistency
시각-언어-행동(VLA) 정책은 로봇 조작 분야에서 뛰어난 성능을 보이지만, 시각 정보, 로봇 상태 및 실행된 행동 간의 시간적 정렬이 깨지는 실시간 교란에 취약합니다. 본 논문에서는 ActFovea라는 플러그 앤 플레이 보호 프레임워크를 소개합니다. ActFovea는 VLA 정책을 재학습하거나 수정하지 않고도 이러한 오류를 감지하고 완화합니다. ActFovea는 로봇 운동학, 고유 센서 데이터 및 최근 행동을 사용하여 접촉과 관련된 영역과 예측된 움직임 경로를 유지하면서 작업과 관련 없는 시각 정보를 억제하는 액션 기반의 주의 집중 영역(foveated regions)을 생성합니다. ActFovea는 시각적 움직임과 관찰 정보의 신선성이 기하학적, 고유 센서 데이터 및 행동 전환과 일관성을 유지하는지 평가하여 실시간 위험을 감지합니다. 복구 가능한 교란의 경우, ActFovea는 교란에 특화된 후보 관찰 데이터를 생성하고, 결과적인 행동 단계를 검증한 후에만 복구를 허용합니다. 오래되거나 재생된 관찰 데이터로 인해 신뢰할 수 있는 복구가 불가능한 경우, 안전한 실패 절차를 실행합니다. 여러 LIBERO 스위트에서 $π_0$ 정책에 대한 폐쇄 루프 평가 결과, ActFovea는 로컬 시각 오버레이 환경에서 성공률을 49.3%에서 90.3%로 향상시켜 깨끗한 성능과의 격차를 93.7% 줄였습니다. 또한 ActFovea는 행동 드리프트 및 시각 지연 환경에서 각각 7.0% 및 9.8%의 성공률을 추가적으로 향상시켰으며, 동시에 깨끗한 작업 환경에서의 성능을 유지했습니다. 관찰 데이터 재생 상황에서는 ActFovea가 모든 실험에서 적시에 안전한 실패를 유도했으며, 보호되지 않은 실패는 발생하지 않았습니다. 이러한 결과는 시공간적 시각-행동 일관성이 VLA 정책의 실시간 보호를 위한 효과적인 기반을 제공한다는 것을 보여줍니다.
Vision-language-action (VLA) policies achieve strong performance in robotic manipulation but remain vulnerable to runtime disturbances that break the temporal alignment among visual observations, robot states, and executed actions. We introduce ActFovea, a plug-and-play safeguarding framework that detects and mitigates such failures without retraining or modifying the underlying VLA policy. ActFovea uses robot kinematics, proprioceptive states, and recent actions to construct action-conditioned foveated regions that retain contact-relevant areas and predicted motion corridors while suppressing task-irrelevant visual content. It detects runtime risks by evaluating whether visual motion and observation freshness remain consistent with geometric, proprioceptive, and action transitions. For recoverable disturbances, ActFovea constructs disturbance-specific candidate observations and accepts a recovery only after verifying the resulting action chunk. When stale or replayed observations make reliable recovery impossible, it invokes a bounded safe-failure procedure. In closed-loop evaluations of $π_0$ across multiple LIBERO suites, ActFovea increases success under localized visual overlays from 49.3\% to 90.3\%, closing 93.7\% of the gap to clean performance. It further improves success under action drift and visual delay by 7.0 and 9.8 percentage points, respectively, while preserving clean-task performance. Under frozen-observation replay, ActFovea triggers timely safe failure in all trials, with no unprotected failures. These results demonstrate that spatiotemporal visual-action consistency provides an effective basis for runtime safeguarding of VLA policies.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.