트랜스포머를 활용한 UTM의 안전 관련 시나리오 분석
Revealing Safety-Critical Scenarios for UTM via Transformer
무인 항공 교통 관리(UTM) 시스템은 여러 대의 항공기를 원격으로 관리하고 조정하도록 설계된 클라우드 기반 플랫폼입니다. UTM 시스템은 충돌이나 추락과 같은 오류를 허용할 수 없는 안전이 중요한 시스템입니다. 잠재적인 취약점을 파악하기 위해, 현재까지 최적의 오류 노출 방법을 제시하거나 명확한 보상 신호를 제공하는 방법은 존재하지 않습니다. 또한, UTM의 자체 복구 기능은 심각한 오류의 '롱테일 효과'를 야기합니다. 본 연구에서는 UTM의 취약점 발견을 시퀀스 모델링 문제로 정의하고, 트랜스포머 기반 강화 학습(RL) 아키텍처를 활용하는 방법을 제안합니다. 저희 방법론은 어텐션 메커니즘을 사용하여 시스템 상태 간의 관계를 직접적으로 모델링하고 최적의 행동을 예측합니다. 저희 프레임워크는 목표 시나리오 생성을 위한 정책 모델과 도메인 제약 조건을 적용하기 위한 액션 샘플러를 도입합니다. 또한, 위험 기반 보상 함수를 사용하여 탐색을 안내합니다. 700시간 규모의 시뮬레이션 연구를 통해, 저희 방법은 전문가가 설계한 테스트에 비해 취약점 발견 효율성을 8배 향상시키는 것을 입증했습니다. 또한, 기존 방법으로는 놓쳤던 중요한 예외적인 사례들을 발견했습니다.
Unmanned Traffic Management (UTM) systems are cloud-based platforms designed to manage and coordinate multiple aerial vehicles remotely. UTM systems are safety-critical which cannot tolerate failures like crash or collision. To reveal latent vulnerabilities, there are neither optimal failure-exposing demonstrations nor clear reward signals. Additionally, UTM's self-healing capability introduces the ``long-tail effect'' of critical failures. We propose framing UTM vulnerability discovery as a sequence modeling problem amenable to transformer-based RL architectures. Our approach leverages attention mechanisms to directly model the relationship among system states, and predict optimal actions. Our framework introduces a Policy Model that generates targeted test scenarios and an Action Sampler that enforces domain constraints. We use a risk-based reward function to guide exploration. Through extensive evaluation on a 700-hour simulation study, we demonstrate an 8$\times$ improvement in vulnerability discovery efficiency compared to expert-guided testing. It also discovers critical edge cases that traditional methods have missed.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.