엔드 투 엔드 UAV 항법을 위한 경량화된 안전 강화 학습
Lightweight Safe Reinforcement Learning for End-to-End UAV Navigation
자율 항공 시스템의 급속한 발전으로 인해 무인항공기(UAV)는 점검, 환경 모니터링 및 구조 등 다양한 분야에 점점 더 많이 활용되고 있으며, 이는 안정적인 자율 항법에 대한 수요를 증가시키고 있습니다. 그러나 빽빽한 환경에서의 자율 UAV 항법은 제한된 인식 능력과 동적 제약 조건 하에서 여전히 어려운 과제입니다. 대부분의 강화 학습(RL) 방법은 명시적인 안전 메커니즘이 부족하여 불안정한 학습, 위험한 행동을 초래하며, 특히 고속 비행 시 더욱 그렇습니다. 또한, 안전 강화 학습 접근 방식에서도 안전성을 확보하기 위해 정책 출력을 안전한 동작 공간으로 투영하는 경우가 많지만, 이로 인해 불안정성이 발생할 수 있습니다. 한편, 많은 머신러닝 기반 방법은 밀집된 입력이나 큰 네트워크를 사용하므로 계산 부담이 증가하고, 경량화된 탑재형 시스템 구현에 제한이 있습니다. 위와 같은 과제에 대응하기 위해, 우리는 UAV 항법을 위한 안전 제약 조건과 인지-제어 통합 프레임워크를 제안합니다. 경량화된 네트워크는 비대칭 및 깊이 분리 컨볼루션을 사용하여 희소한 관측 데이터를 충돌 위험을 고려하는 특징으로 변환합니다. 우리는 계층적 제어 구조 내에서 이 문제를 제한된 마르코프 의사 결정 프로세스로 정의하고, Lagrangian 기반의 안전 PPO 알고리즘을 사용하여 해결합니다. 커리큘럼 학습은 추가적으로 학습 안정성을 향상시킵니다. 다양한 장애물 밀도와 비행 속도로 진행한 실험 결과, 제안하는 방법은 기존 강화 학습 기준보다 높은 성공률, 향상된 안전성 및 더 나은 효율성을 보여주었습니다.
With the rapid development of autonomous aerial systems, Unmanned Aerial Vehicles (UAVs) are increasingly deployed in applications such as inspection, environmental monitoring, and rescue, creating growing demand for reliable autonomous navigation. However, autonomous UAV navigation in dense environments remains challenging under sparse perception and dynamic constraints. Most reinforcement learning (RL) methods lack explicit safety mechanisms, leading to unsafe exploration, unstable training, and risky behaviors, especially during high-speed flight. Even in safe RL approaches, safety is often enforced by projecting policy outputs onto a safe action set, which may introduce instability. Meanwhile, many learning-based methods rely on dense inputs or large networks, increasing computational burden and limiting lightweight onboard deployment. Facing the above challenges, we propose a safety-constrained perception-control integrated framework for UAV navigation. A lightweight network encodes sparse observations into collision-risk-aware features using asymmetric and depthwise separable convolutions. We formulate the task as a constrained Markov decision process within a hierarchical control architecture and solve it using a Lagrangian-based safe PPO algorithm. Curriculum learning further improves training stability. Experiments with varying obstacle densities and flight speeds demonstrate higher success rates, improved safety, and better efficiency than existing reinforcement learning baselines.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.