2603.07973v1 Mar 09, 2026 cs.RO

VORL-EXPLORE: 동적 환경에서의 다중 로봇 탐색을 위한 하이브리드 학습 기반 계획 접근 방식

VORL-EXPLORE: A Hybrid Learning Planning Approach to Multi-Robot Exploration in Dynamic Environments

Shangke Lyu
Shangke Lyu
Citations: 534
h-index: 12
Ninghao Liu
Ninghao Liu
Citations: 0
h-index: 0
Sen Shen
Sen Shen
Citations: 4
h-index: 1
Zheng Li
Zheng Li
Citations: 27
h-index: 4
Sheng Liu
Sheng Liu
Karlsruhe Institute of Technology
Citations: 83
h-index: 3
Dongkun Han
Dongkun Han
Citations: 4
h-index: 1
T. Braunl
T. Braunl
Citations: 297
h-index: 6

계층적 다중 로봇 탐색은 일반적으로 전선 할당과 로컬 탐색을 분리하여 수행하는데, 이는 밀집되고 동적인 환경에서 시스템의 안정성을 저해할 수 있습니다. 전선 할당기는 실행 난이도에 대한 직접적인 정보를 갖지 못하기 때문에, 로봇들이 병목 지점에 뭉치거나, 진동적인 재계획을 유발하거나, 중복적인 탐색을 수행할 수 있습니다. 본 논문에서는 VORL-EXPLORE라는 하이브리드 학습 및 계획 프레임워크를 제안합니다. 이 프레임워크는 '실행 충실도(execution fidelity)'라는 개념을 통해 이러한 한계를 극복합니다. 실행 충실도는 로컬 탐색 가능성에 대한 공유된 추정치로서, 작업 할당과 움직임 실행을 결합합니다. 이 충실도 신호는 로봇 간의 반발력을 고려한 충실도 기반 보로노이 목표 함수에 통합되어, 문제가 발생하기 전에 충돌 가능성을 줄입니다. 또한, 이 신호는 전역 A* 경로 탐색과 반응형 강화 학습 정책 간의 위험 기반 적응형 조정 메커니즘을 구동하여, 장거리 효율성과 좁은 공간에서의 안전한 상호 작용을 균형 있게 유지합니다. 또한, 본 프레임워크는 최근 진행 상황과 안전 결과를 기반으로 생성된 가짜 레이블을 사용하여 충실도 모델을 온라인으로 자체적으로 재보정할 수 있는 기능을 지원합니다. 이를 통해 수동적인 위험 설정 없이 비정상적인 장애물에 대한 적응이 가능합니다. 제안하는 프레임워크의 성능을 평가하기 위해, 교통 체증 상황에 대한 별도의 실험을 진행했습니다. 무작위 그리드 환경 및 Gazebo 공장 시뮬레이션 환경에서의 광범위한 실험 결과, 높은 성공률, 짧은 경로 길이, 낮은 중복, 그리고 강력한 충돌 회피 성능을 확인했습니다. 본 논문이 채택되면 소스 코드를 공개할 예정입니다.

Original Abstract

Hierarchical multi-robot exploration commonly decouples frontier allocation from local navigation, which can make the system brittle in dense and dynamic environments. Because the allocator lacks direct awareness of execution difficulty, robots may cluster at bottlenecks, trigger oscillatory replanning, and generate redundant coverage. We propose VORL-EXPLORE, a hybrid learning and planning framework that addresses this limitation through execution fidelity, a shared estimate of local navigability that couples task allocation with motion execution. This fidelity signal is incorporated into a fidelity-coupled Voronoi objective with inter-robot repulsion to reduce contention before it emerges. It also drives a risk-aware adaptive arbitration mechanism between global A* guidance and a reactive reinforcement learning policy, balancing long-range efficiency with safe interaction in confined spaces. The framework further supports online self-supervised recalibration of the fidelity model using pseudo-labels derived from recent progress and safety outcomes, enabling adaptation to non-stationary obstacles without manual risk tuning. We evaluate this capability separately in a dedicated severe-traffic ablation. Extensive experiments in randomized grids and a Gazebo factory scenario show high success rates, shorter path length, lower overlap, and robust collision avoidance. The source code will be made publicly available upon acceptance.

1 Citations
0 Influential
6 Altmetric
31.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!