구조 인식 기반의 견고한 미세 조정: 물리적 주의 우회 공격으로부터 비전-언어-행동 로봇을 보호
Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking
비전-언어-행동(VLA) 정책은 일반적인 로봇 조작을 가능하게 하지만, 실제 환경에서의 공격에 대한 안정성은 여전히 취약합니다. 본 연구에서는 물리적으로 구현 가능한 적대적 패치가 특정 메커니즘을 활성화하여 오류를 유발할 수 있음을 보여줍니다. 이 메커니즘은 '정책-중요한 행동-to-비전 주의 우회'로, 여기서 행동에 따른 주의가 작업과 관련된 영역에서 벗어나 국소적인 패치로 전환됩니다. 이러한 위협을 입증하기 위해, 본 연구에서는 Attention-Guided Semantic Disruption (AGSD)를 제안합니다. AGSD는 Expectation-over-Transformation (EOT)으로 최적화된 인쇄 가능한 패치이며, (i) 행동에 따른 주의를 패치로 집중시키고 (ii) 비전-언어 의미 정렬을 방해하여 강력한 교차 작업 및 교차 아키텍처 성능을 달성합니다. 이러한 공격을 완화하기 위해, 구조 인식 기반의 견고한 미세 조정(Structure-Aware Robust Fine-Tuning, SARF)이라는 새로운 방어 기법을 제안합니다. SARF는 특징 고정, 정책-중요한 주의 보정 및 언어를 활용한 기하학적 일관성 제약을 통해 시각 인코더만 미세 조정하여 0%의 추론 오버헤드로 작동합니다. LIBERO 데이터셋에서 SARF는 AGSD 공격 하에 OpenVLA의 실패율을 100%에서 28.6% (평균)로 감소시키면서, 정상적인 성능은 유지했습니다. 또한 실제 PiPER 조작기를 사용하여 평균 성공률을 23.0%에서 65.0%로 향상시켰습니다. 이러한 결과는 메커니즘 수준의 안정성이 물리적 주의 우회 공격으로부터 VLA 로봇을 보호하는 실용적인 방법임을 보여줍니다.
Vision-Language-Action (VLA) policies promise general robotic manipulation, but their robustness against physical-world attacks remains fragile. In particular, we show that physically realizable adversarial patches can reliably induce failures by triggering a mechanism we call policy-critical action-to-vision attention hijacking, where action-conditioned attention is diverted from task-relevant regions to a localized patch. To demonstrate the threat, we propose Attention-Guided Semantic Disruption (AGSD), an Expectation-over-Transformation (EOT) optimized printable patch that jointly (i) concentrates action-to-vision attention on the patch and (ii) disrupts vision-language semantic alignment, yielding strong cross-task and cross-architecture transfer. To mitigate such attacks, we introduce Structure-Aware Robust Fine-Tuning (SARF), a zero-inference-overhead defense that fine-tunes only the visual encoder using feature anchoring, policy-critical attention correction, and language-guided geometric consistency restricted to semantically relevant regions. On LIBERO, SARF reduces OpenVLA's failure rate under AGSD from 100% to 14.2%-56.8% (28.6% average) across suites while preserving clean performance, and on a real PiPER manipulator it improves average success under AGSD from 23.0% to 65.0%. These results highlight mechanism-level robustness as a practical path to securing VLA robots against physical attention hijacking.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.