AdvNav: 행동 기반의 블랙박스 적대적 공격 - 시각-언어 내비게이션 시스템을 대상으로
AdvNav: Behavior-Guided Black-Box Adversarial Attacks on Vision-Language Navigation
임베디드 AI 분야에서 상당한 발전이 있었음에도 불구하고, 시각-언어 내비게이션(VLN) 시스템은 여전히 적대적인 시각적 교란에 취약합니다. 대부분의 기존 방법은 대상 모델의 기울기 정보 접근을 필요로 하지만, 이는 실제 환경에 배포된 시스템에서는 비현실적이며, 최적화를 위한 재귀적 역전파 과정 때문에 계산적으로 매우 부담스럽습니다. 또한, 이전의 블랙박스 공격 방법들은 주로 단일 단계, 즉각적인 의사 결정 작업을 대상으로 하며, 복잡한 작업과 시간적 의존성을 처리하는 데 어려움을 겪습니다. 따라서, 관측 가능한 입력 및 출력만 사용하여 다단계 순차적 인지-행동 루프를 효과적으로 방해할 수 있는 기울기 기반 공격이 아닌 방법을 개발해야 합니다. 이에 우리는 AdvNav라는 행동 기반의 블랙박스 적대적 공격 프레임워크를 제안합니다. 이 프레임워크는 내비게이션 과정에서 에이전트의 시점에서 보이는 이미지에 교란을 가합니다. 블랙박스 환경에서 효과적인 최적화를 위한 정보 제공적인 대리 목적 함수(surrogate objective)를 구축하기 위해, 우리는 2가지 수준의 행동 기반 피드백을 설계했습니다. 여기에는 전체 내비게이션 성능 저하를 나타내는 경로 레벨 성과 점수, 잠재적인 의사 결정 위험을 고려하는 행동 레벨 보상 점수, 그리고 에이전트 자체의 행동으로부터 추출된 편차 지표가 포함됩니다. 이 피드백은 하이브리드 최적화 전략을 안내합니다. 이 전략은 적응형 업데이트를 통해 교란 강도를 휴리스틱하게 조정하고, 유전 알고리즘을 사용하여 노이즈의 공간 구조를 진화시켜 반복적으로 가장 파괴적인 노이즈 구성을 찾습니다. R2R 데이터셋에서 Transformer 기반 HAMT와 LLM 기반 MapGPT의 두 가지 아키텍처에 대해 AdvNav를 평가한 결과, 공격 성공률은 각각 49.70%, 65.96%, 87.30%로 나타났습니다. 이러한 결과는 AdvNav의 효과성과 일반성을 입증하고, VLN 시스템의 중요한 인지 취약점을 드러내며, 향후 보다 안정적인 VLN 모델 설계에 대한 통찰력을 제공합니다.
Despite progress in Embodied AI, Vision-and-Language Navigation systems remain vulnerable to adversarial visual disturbances. Most existing methods rely on white-box access to target model gradients, which is often unrealistic for real-world deployed systems and computationally exhaustive due to recursive backpropagation for optimization, limiting their applicability. While previous black-box methods predominantly target single-step, instantaneous decision tasks, they struggle to handle the task complexities and temporal dependencies. This highlights the need for a gradient-free attack method that can effectively disrupt the multistep sequential perception-action loop using only observable inputs and outputs. Therefore, we propose AdvNav, a behavior-guided black-box adversarial attack framework that disturbs an agent's first-person views during navigation. To construct an informative surrogate objective for effective optimization guidance in gradient-free search under the black-box setting, we design a dual-granularity behavior-based feedback, aggregating a trajectory-level performance score representing overall navigation degradation, an action-level reward score considering the potential decision risk, and a deviation indicator, all of which are extracted from the agent's self-output behaviors. This feedback guides a hybrid optimization strategy that heuristically tunes perturbation strength via adaptive updates and evolves noise spatial structure genetically, to iteratively discover the most disruptive noise configuration. Evaluated against Transformer-based HAMT and LLM-based MapGPT with two types of backbones on R2R dataset, AdvNav achieves 49.70/65.96/87.30% Attack Success Rate. The result demonstrates the effectiveness and generality of AdvNav, reveals critical perception vulnerabilities and offers insights for the design of future resilient VLN models.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.