2607.29009v1 Jul 31, 2026 cs.RO

D-VLC: 미지의 환경에서 다양한 유형의 로봇 시스템을 위한 분산형 시각-언어 협업

D-VLC: Decentralized Vision-Language Collaboration for Heterogeneous Embodied Multi-Robot Systems in Unknown Environments

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Fei Gao
Fei Gao
Citations: 77
h-index: 3
Yuan Zhou
Yuan Zhou
Citations: 60
h-index: 2
Ruitong Lin
Ruitong Lin
Citations: 0
h-index: 0
Shengyu Wang
Shengyu Wang
Citations: 0
h-index: 0
Weiqi Gai
Weiqi Gai
Citations: 117
h-index: 3
Yuze Wu
Yuze Wu
Citations: 216
h-index: 7

다중 로봇 시스템, 특히 이질적인 로봇 군집은 병렬 협력과 상호 보완적인 기능을 통해 복잡한 작업 수행 효율성을 향상시킬 수 있습니다. 그러나 기존의 규칙 기반 방법은 미리 정의된 작업 모델 및 특수 의사 결정 프로그램에 의존하기 때문에 복잡한 의미 지침을 이해하고 이질적인 로봇을 조정하는 데 어려움이 있습니다. LLM(대규모 언어 모델)은 강력한 언어 이해 및 작업 추론 능력을 제공하여 다중 로봇 시스템이 지침을 해석하고, 작업을 분해하며, 작업 의미에 따라 역할을 할당할 수 있도록 합니다. VLM(시각-언어 모델)은 시각적 인지 기능을 추가하여 로봇이 물리적 환경에서 객체, 영역 및 공간 관계에 대해 추론할 수 있도록 합니다. 그러나 기존의 LLM/VLM 기반 방법은 종종 알려진 지도, 중앙 집중식 및 동기화된 의사 결정에 의존하며, 이는 이질적인 로봇과 새로운 작업으로의 일반화 능력을 제한합니다. 따라서 우리는 분산형 비동기 추론, 경량 정보 공유, 능력 기반 협력 및 통합 액션 인터페이스를 결합하는 프레임워크를 제안합니다. 이를 통해 범용 VLM이 특정 로봇에 대한 작업을 생성하고, 작업 또는 로봇별 훈련 없이 학습된 전문가를 통해 해당 작업을 실행할 수 있습니다. 다양한 시나리오와 여러 VLM을 사용한 실험 결과, 성공률은 70% 이상으로 나타났으며, 완료 시간은 기하학적 탐색 기반(geometric greedy baseline) 방법에 비해 최대 55.8% 단축되었습니다.

Original Abstract

Multi-robot systems, particularly heterogeneous robot swarms, can improve the efficiency of complex task execution through parallel collaboration and complementary capabilities. However, conventional rule-based methods rely on predefined task models and specialized decision making programs, making it difficult to understand complex semantic instructions and coordinate heterogeneous robots. LLMs introduce strong language understanding and task reasoning capabilities, allowing multi-robot systems to interpret instructions, decompose tasks, and assign roles according to task semantics. VLMs further incorporate visual perception, enabling robots to reason about objects, regions, and spatial relationships in physical environments. Nevertheless, existing LLM/VLM based methods often depend on known maps, centralized and synchronized decision making, limiting their generalization to heterogeneous robots and unseen tasks. We therefore propose a framework that combines decentralized asynchronous reasoning, lightweight information sharing, capability aware collaboration, and a unified action interface, enabling general purpose VLMs to generate robot specific actions executed by learning free experts without task or robot specific training. Experiments across diverse scenarios and multiple VLMs show success rates above 70\%, with completion time reduced by up to 55.8\% relative to the geometric greedy baseline.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!