CVPR 2026@AdvML 워크샵 챌린지 기술 보고서
Technical Report on the CVPR 2026@AdvML Workshop Challenge
비전-언어 에이전트(VLA)는 복잡한 운전 장면을 해석하고 안전과 관련된 추론을 지원하는 데 점점 더 많이 사용되고 있습니다. 본 보고서는 자율 주행 VLA에 대한 적대적 다중 모드 공격에 관한 CVPR 2026@AdvML 워크샵 챌린지를 소개합니다. 이 챌린지는 DriveLM 스타일의 멀티뷰 시각 질의 응답을 기반으로 하며, 각 장면을 여섯 개의 동기화된 카메라 이미지와 운전 관련 질문-응답 쌍의 구조화된 컬렉션으로 표현합니다. 참가자들은 모델의 응답이 기준 답변과 다르게 나타나도록 유도하는 동시에 이미지 품질을 유지하고 텍스트 비용을 제한하는 적대적 이미지를 생성하고, 텍스트에 대한 특정 수정을 수행합니다. 이 대회는 두 단계로 구성되며, II단계에서는 전송 가능성을 평가하기 위해 숨겨진 블랙박스 모델이 추가됩니다. 본 보고서에서는 과제 설계, 제출 규칙, 평가 프로토콜 및 리더보드 결과를 설명하고, 기술 보고서가 제공된 다섯 가지 주요 제출 결과를 분석합니다. 이러한 보고서를 통해 다음과 같은 몇 가지 일반적인 패턴이 나타났습니다: 이미지 측 공격은 접미사 페널티에 의해 선호되는 경향이 있으며, 장면 수준의 멀티뷰 최적화는 개별 뷰를 처리하는 것보다 효과적입니다; 질문-응답 유형 및 그래프 구조는 공격 예산을 할당하는 데 유용한 사전 정보를 제공하며, 특징 공간 목표는 블랙박스 전송을 개선할 수 있습니다. 또한, 카메라 이미지에 포함된 타자 콘텐츠는 자율 주행 VLA에서 지속적인 취약점을 드러냅니다. 이러한 결과는 다중 모드 자율 주행 시스템의 향후 견고성 평가 및 방어 설계에 대한 실질적인 참고 자료를 제공합니다.
Vision-language agents (VLAs) are increasingly used to interpret complex driving scenes and support safety-critical reasoning. This report presents the CVPR 2026@AdvML Workshop Challenge on adversarial multimodal attacks against autonomous-driving VLAs. Built on DriveLM-style multi-view visual question answering, the challenge represents each scene with six synchronized camera images and a structured collection of driving-related question-answer pairs. Participants generate adversarial images and suffix-only textual perturbations that induce model responses to deviate from reference answers while preserving image fidelity and limiting textual cost. The competition comprises two phases, with Phase II adding a hidden black-box model to assess transferability. We describe the task design, submission rules, evaluation protocol, and leaderboard results, and then examine five leading submissions for which technical reports were available. Across these reports, several recurring patterns emerge: image-side attacks are favored by the suffix penalty; scene-level, multi-view optimization is more effective than treating views in isolation; QA types and graph structure provide useful priors for allocating attack budget; feature-space objectives can improve black-box transfer; and typographic content embedded in camera images exposes a persistent vulnerability in driving VLAs. These findings provide a practical reference for future robustness evaluation and defense design in multimodal autonomous-driving systems.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.