CoNav-UAV: 스태클버그 학습을 통한 협력적 양고도 항공 네비게이션
CoNav-UAV: Cooperative Dual-Altitude Aerial Navigation via Stackelberg Learning
재난 구조, 기반 시설 점검 및 보안 순찰과 같은 임무를 위해 공중 플랫폼에서의 목표 지향형 시각-언어 내비게이션(VLN) 연구가 활발하게 진행되고 있습니다. 이 과제에서 무인 항공기(UAV)는 대상의 외관과 주변 환경에 대한 간결한 설명만으로 대상을 찾아야 합니다. 이는 전역 탐색 및 객체 인식 능력뿐만 아니라 충돌을 피하는 정밀 접근 능력을 요구하며, 이 두 가지 과정은 단일 에이전트 내에서 조화시키기 어렵습니다. 대부분의 기존 방법은 지상 VLN 패러다임을 저고도 UAV에 적용하고 외부 지원으로 비효율적인 탐색 문제를 보완합니다. 최근에는 상호 보완적인 고도에 두 대의 UAV를 배치하는 시도가 있었지만, 여전히 특권 정보를 사용하며 두 에이전트를 독립적으로 훈련하여 협력에 필수적인 상호 적응을 방해합니다. 본 연구에서는 시스템이 온보드 시각 및 언어 입력만으로 작동하도록, 고고도 리더와 저고도 팔로워 간의 스태클버그 게임으로 과제를 명시적으로 모델링하는 CoNav-UAV를 제안합니다. 이 게임을 해결하기 위해 Iterative Stackelberg Learning 방법을 도입했습니다. 리더의 고수준 시각-언어 추론은 메모리 기반의 문맥 내 학습을 통해 개선되고, 팔로워의 정밀한 동작 제어는 DAgger 방식의 전문가 지식 전달을 통해 업데이트됩니다. 이러한 반복적인 과정은 두 에이전트를 스태클버그 균형 상태로 이끌도록 합니다. CoNav-UAV는 AerialVLN 벤치마크에서 제공하는 세 개의 고품질 도시 환경에서 단일 에이전트 및 이중 에이전트 기반 모델보다 우수한 성능을 보였습니다. 학습 환경에서 성공률이 최대 30.8% 향상되었으며, 교차 장면 전송 시에는 9.0% 향상되었습니다. 또한, 약 3배 적은 양의 적응 데이터만 사용했습니다. 추가 분석 결과, 리더 및 팔로워 업데이트가 상호 보완적인 이점을 제공하며, VLM 백본에 따라 학습 동역학이 다르지만 강력한 성능 향상을 가져옴을 확인했습니다.
Target-oriented vision-and-language navigation (VLN) on aerial platforms is attracting growing attention for missions such as disaster rescue, infrastructure inspection, and security patrol. In this task, an unmanned aerial vehicle (UAV) needs to locate targets given only a concise description of their appearance and surroundings. This requires global exploration and grounding as well as collision-free close-range approach, two interleaved processes difficult to reconcile within a single agent. Most existing methods transfer the ground VLN paradigm to a low-altitude UAV and compensate for its inefficient exploration with external assistance. A recent attempt deploys two UAVs at complementary altitudes yet still relies on privileged information and trains its two agents independently, precluding any mutual adaptation essential for cooperation. Here we propose CoNav-UAV, which explicitly models the task as a Stackelberg game between a high-altitude leader and a low-altitude follower, with the system operating on onboard visual and linguistic inputs alone. To solve this game, we introduce Iterative Stackelberg Learning. The leader's high-level vision-language reasoning is refined via memory-based in-context learning, while the follower's precise motion control is updated via DAgger-style expert distillation. The alternation drives both agents toward a Stackelberg equilibrium. CoNav-UAV consistently outperforms single- and dual-agent baselines across three high-fidelity urban scenes from the AerialVLN benchmark. Success rate improves by up to 30.8 points on the learning scene, and 9.0 points under cross-scene transfer while using about 3x less adaptation data. Further analyses validate the complementary gains of the leader and follower updates and reveal robust gains yet distinct learning dynamics across VLM backbones.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.