2603.15221v1 Mar 16, 2026 cs.LG

ADV-0: 자율 주행 시스템의 롱테일 강건성을 위한 폐루프 민-맥 적대적 학습

ADV-0: Closed-Loop Min-Max Adversarial Training for Long-Tail Robustness in Autonomous Driving

Tong Nie
Tong Nie
Citations: 4
h-index: 1
Yihong Tang
Yihong Tang
Citations: 110
h-index: 3
Junlin He
Junlin He
Citations: 104
h-index: 6
Yuewen Mei
Yuewen Mei
Citations: 223
h-index: 7
Lijun Sun
Lijun Sun
Citations: 15
h-index: 2
Jian Sun
Jian Sun
Citations: 24
h-index: 3
Jie Sun
Jie Sun
Citations: 76
h-index: 6
Wei Ma
Wei Ma
Citations: 261
h-index: 7

자율 주행 시스템의 배포에는 드물지만 안전에 매우 중요한 롱테일 시나리오에 대한 강건성이 요구됩니다. 적대적 학습은 유망한 해결책을 제공하지만, 기존 방법은 일반적으로 시나리오 생성과 정책 최적화를 분리하고 휴리스틱적인 대리 방법을 사용합니다. 이는 목표 불일치를 초래하며, 진화하는 정책의 변화하는 실패 모드를 포착하지 못합니다. 본 논문에서는 ADV-0이라는 폐루프 민-맥 최적화 프레임워크를 제시합니다. 이 프레임워크는 자율 주행 정책(방어자)과 적대적 에이전트(공격자) 간의 상호 작용을 영-합 마르코프 게임으로 간주합니다. 공격자의 유틸리티를 방어자의 목표와 직접적으로 일치시켜 최적의 적대적 분포를 파악합니다. 이를 실용적으로 구현하기 위해, 동적인 적대적 진화를 반복적인 선호도 학습으로 표현하여 최적값을 효율적으로 근사하고, 알고리즘에 독립적인 솔루션을 제공합니다. 이론적으로 ADV-0는 내쉬 균형으로 수렴하고, 실제 성능의 인증된 하한을 최대화합니다. 실험 결과, ADV-0는 다양한 안전에 중요한 실패를 효과적으로 드러내고, 학습된 정책과 모션 플래너의 일반화 성능을 향상시켜 예측하지 못한 롱테일 위험에 대한 강건성을 크게 향상시킵니다.

Original Abstract

Deploying autonomous driving systems requires robustness against long-tail scenarios that are rare but safety-critical. While adversarial training offers a promising solution, existing methods typically decouple scenario generation from policy optimization and rely on heuristic surrogates. This leads to objective misalignment and fails to capture the shifting failure modes of evolving policies. This paper presents ADV-0, a closed-loop min-max optimization framework that treats the interaction between driving policy (defender) and adversarial agent (attacker) as a zero-sum Markov game. By aligning the attacker's utility directly with the defender's objective, we reveal the optimal adversary distribution. To make this tractable, we cast dynamic adversary evolution as iterative preference learning, efficiently approximating this optimum and offering an algorithm-agnostic solution to the game. Theoretically, ADV-0 converges to a Nash Equilibrium and maximizes a certified lower bound on real-world performance. Experiments indicate that it effectively exposes diverse safety-critical failures and greatly enhances the generalizability of both learned policies and motion planners against unseen long-tail risks.

2 Citations
0 Influential
3.5 Altmetric
19.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!