비행 전에 신중하게: 시각 정보를 활용한 공간적 판단을 통한 드론의 목표 접근 및 정지
Deliberate Before You Fly: Vision-Guided Spatial Deliberation for UAV See-and-Reach Navigation
드론이 초기 시야에서 보이는 언어적으로 지정된 목표 지점에 접근하고 안정적으로 근처에 멈추는 '시각-접근' 방식의 탐색은, 기존 방법들이 일반적으로 시각-언어 정보를 직접적인 행동 결과로 변환하는 반면, 중간 단계의 세밀한 공간적 판단을 명시적으로 모델링하지 않아 의미와 제어 간의 불일치가 발생하고, 이는 일관성 없는 움직임과 신뢰할 수 없는 종료 지점으로 이어질 수 있습니다. 이러한 문제를 해결하기 위해, 우리는 시각 정보를 활용하여 목표 지점 경로 생성을 위한 공간적 판단 단계를 도입하는 'DBFly'라는 프레임워크를 제안합니다. 구체적으로, DBFly는 목표 방향을 기준으로 앵커링, 공간 진단 및 기동 결정으로 구성된 일련의 과정을 통해 고수준의 기동 의도를 명시적으로 반영하여 연속적인 경로 생성을 가능하게 합니다. 또한, 초기 목표 방향 정보를 지속적인 기하학적 참조점으로 변환하고 드론의 현재 위치에서 실시간으로 해당 영역 상태를 추정함으로써, 공간 진단 및 기동 수정에 대한 부드러운 기하학적 지침을 제공하는 숨겨진 비행 경로를 생성합니다. 더불어, DBFly는 목표와의 근접성뿐만 아니라 단기적인 움직임 수렴을 통해 종료 상태를 정의하는 전략을 사용하여 목표 지점 주변에서 더욱 안정적으로 멈추도록 합니다. 다양한 환경에서의 실험 결과, DBFly는 기존 최고 성능 모델보다 평균 25.07%의 성공률 향상을 보여주었습니다. 프로젝트 홈페이지는 https://xuefanfu.github.io/DBFly-Page 에서 확인할 수 있습니다.
UAV see-and-reach navigation requires an aerial agent to approach a language-specified target visible in its initial view and stop reliably near it. Existing methods typically map vision-language representations directly to action outputs without explicitly modeling intermediate fine-grained spatial decisions. This direct mapping causes semantic-control misalignment, leading to inconsistent maneuvers and unreliable termination. To address this issue, we propose DBFly, a vision-language waypoint prediction framework that introduces explicit vision-guided spatial deliberation before waypoint generation. Specifically, DBFly introduces a spatial maneuver decision chain that progressively performs target-direction anchoring, spatial diagnosis, and maneuver decision, enabling high-level maneuver intent to explicitly guide continuous waypoint generation. DBFly further constructs an implicit flight corridor by transforming the initial target-direction prior into a persistent geometric reference and deriving an online corridor state from the UAV's current position, thereby providing soft geometric guidance for spatial diagnosis and maneuver correction. In addition, DBFly develops a terminal-convergence-aware stopping strategy that characterizes terminal states through both target proximity and short-horizon motion convergence, enabling more reliable stopping near the target. Extensive experiments across seen, unseen-object, and unseen-scene test sets demonstrate that DBFly improves the success rate over the SOTA baseline by an average of 25.07 percentage points. The project homepage is available at https://xuefanfu.github.io/DBFly-Page.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.