시야 범위 내 UAV를 위한 정밀 시각-언어 기반 네비게이션: See-and-Reach
See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View
UAV 비전-언어 네비게이션(UAV-VLN)은 일반적으로 장거리 목표 탐색과 최종 목표 접근을 동시에 최적화하고 평가하는 전체적인 검색 및 도달 문제로 정의됩니다. 이러한 방식은 항공 로봇의 중요한 능력, 즉 UAV가 보이는 목표물을 정확하게 인식하고 목표물이 시야 범위 내에 들어오면 시각-언어 정보를 바탕으로 정밀한 3차원 움직임을 생성할 수 있는지 여부를 평가하기 어렵다는 한계를 가지고 있습니다. 이 문제를 해결하기 위해, 우리는 시야 범위 내에서 목표물을 직접적으로 탐색하는 See-and-Reach 단계를 분리하고 최종 도달 능력을 보다 정확하게 평가할 수 있는 UAV-VLN-FOV라는 새로운 네비게이션 태스크를 제안합니다. 또한, 미세한 시각적 정보 인식과 공간 방향 일관성을 향상시켜 정밀한 목표 도달을 가능하게 하는 3DG-VLN이라는 비전-언어 웨이포인트 예측 프레임워크를 개발했습니다. 구체적으로, 3DG-VLN은 고해상도 전방 및 하향 시야 관찰 데이터를 적응적으로 처리하여 목표물 인식에 필요한 미세한 시각적 및 기하학적 정보를 유지합니다. 또한, 클로즈드 루프 네비게이션 과정에서 목표물을 기준으로 하는 방향을 실시간으로 업데이트하여 에이전트가 목표물과의 공간적 정렬 상태를 유지하고 누적되는 방향 오차를 줄일 수 있도록 합니다. 이 태스크를 지원하기 위해, 2,717개의 트랙케이스로 구성된 고해상도 벤치마크를 구축했습니다. 이 벤치마크는 목표 지향적인 상위 레벨 명령어, 고해상도 전방 및 하향 시야의 자기 중심 관찰 데이터, 그리고 연속적인 3차원 웨이포인트 주석을 포함합니다. 실험 결과, 3DG-VLN은 기존의 UAV-VLN 모델보다 우수한 성능을 보이며 성공률이 13.82% 향상되었습니다. 실제 환경에서의 테스트 결과는 3DG-VLN이 실용적인 See-and-Reach 네비게이션에 잠재력을 가지고 있음을 보여줍니다. 소스 코드 및 벤치마크는 https://github.com/xuefanfu/3DG-VLN 에서 확인할 수 있습니다.
UAV Vision-Language Navigation (UAV-VLN) is typically formulated as a holistic search-and-reach problem, where long-range target discovery and final target approach are optimized and evaluated jointly. This formulation makes it difficult to assess a critical capability of aerial embodied agents, namely whether a UAV can accurately ground a visible target and translate vision-language evidence into precise 3D motion once the target enters its field of view. To address this limitation, we introduce UAV-VLN-FOV, a target-visible navigation task that isolates the see-and-reach stage and enables a more diagnostic evaluation of terminal reaching ability. We further propose 3DG-VLN, a vision-language waypoint prediction framework guided by dynamic 3D direction cues to enhance fine-grained visual grounding and spatial direction alignment for precise target reaching. Specifically, 3DG-VLN adaptively processes high-resolution front-view and downward-view observations to preserve fine-grained visual and geometric details for target grounding. It also updates the target-relative direction online during closed-loop navigation, allowing the agent to maintain spatial alignment with the target and reduce accumulated direction drift. To support this task, we construct a dedicated high-resolution benchmark which contains 2,717 trajectories with target-oriented high-level instructions, high-resolution front-view and downward-view egocentric observations, and continuous 3D waypoint annotations. Experiments show that 3DG-VLN outperforms competitive UAV-VLN baselines, achieving a 13.82\% improvement in success rate. Real-world trials further demonstrate the potential of 3DG-VLN for practical see-and-reach navigation. The source code and benchmark are available at https://github.com/xuefanfu/3DG-VLN.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.