2603.07822v2 Mar 08, 2026 cs.RO

언제 질문해야 할까: 명시적인 대화와 암묵적인 의도 단서를 활용하여 인간-로봇 공동 계획의 불확실성 해결

Knowing When to Ask: Resolving Uncertainty in Human-Robot Joint Planning via Explicit Dialogue and Implicit Intent Cues

Tian Lan
Tian Lan
Citations: 28
h-index: 3
Zeyu Fang
Zeyu Fang
Citations: 48
h-index: 5
Mahdi Imani
Mahdi Imani
Citations: 107
h-index: 6
Rongqian Chen
Rongqian Chen
Citations: 25
h-index: 3
Beomyeol Yu
Beomyeol Yu
Citations: 157
h-index: 7
Yuxin Lin
Yuxin Lin
Citations: 41
h-index: 4
Chenghao Liu
Chenghao Liu
Citations: 10
h-index: 2
Zeyuan Yang
Zeyuan Yang
Citations: 158
h-index: 4
Taeyoung Lee
Taeyoung Lee
Citations: 27
h-index: 2

개방형 환경에서 효과적인 인간-로봇 협업은 작업, 환경 및 인간 동료에 대한 불확실성이 존재하는 상황에서의 공동 계획을 필요로 합니다. 의사소통은 이러한 불확실성을 해소하는 가장 직접적인 방법이지만, 대부분의 기존 시스템은 일방향 통신만을 지원합니다. 즉, 로봇은 듣고 행동하며, 인간을 능동적인 감독자가 아닌 양방향 대화가 가능한 협력 파트너로 여기지 않습니다. 본 연구에서는 로봇이 두 가지 상호 보완적인 의사소통 채널을 통해 불확실성을 적극적으로 해결하는 통합된 인간-로봇 공동 계획 시스템을 제안합니다. 의사 결정에 중요한 불확실성이 존재할 경우, 불확실성 완화 공동 계획 모듈은 인간과의 명확화 대화를 통해 작업 지시의 모호성을 LLM 기반의 능동적인 정보 획득 메커니즘을 사용하여 해결하고, 가능 경로 가설을 A* 탐색 알고리즘에 적용하며, 동적 프로그래밍을 통해 최적 비용의 질문 정책을 계산하여 로봇이 계획에 실제로 중요한 질문만을 하도록 합니다. 명시적인 대화가 불필요하거나 실용적이지 않은 경우, 실시간 의도 인지 협업 모듈은 암묵적인 비언어적 단서를 해석하여 공간 및 방향 신호를 기반으로 인간의 잠재적인 작업 의도를 확률적으로 추론하고, 통신 오버헤드 없이 협력에 적합한 작업을 선택할 수 있도록 합니다. 제안된 시스템은 Gazebo 시뮬레이션과 실제 UAV 환경에서 검증되었으며, 음성 대화 인터페이스와 Vision-Language Model (VLM) 기반의 3D 의미 인식 파이프라인과 통합되었습니다. 실험 결과는 최적 비용의 명확화 대화가 상호 작용 비용을 51.9% 절감하면서도 100%의 작업 성공률을 유지하며, 암묵적인 의도 해석은 기존 방식에 비해 협업 작업 실행 시간을 25.4% 단축한다는 것을 보여줍니다.

Original Abstract

Effective human-robot collaboration in open-world environments requires joint planning under uncertainty about the task, the environment, and the human teammate. Communication is the most direct means of resolving such uncertainty, yet most existing systems support only one-way communication: robots listen and act, treating humans as passive supervisors rather than conversational teammates capable of two-way dialogue. We propose a unified human-robot joint planning system in which the robot actively resolves uncertainty through two complementary communication channels. When uncertainty is decision-critical, an uncertainty-mitigation joint planning module engages the human in clarification dialogue: it grounds ambiguous instructions via an LLM-assisted active elicitation mechanism, enumerates traversability hypotheses through a hypothesis-augmented A* search, and computes a cost-optimal querying policy via dynamic programming, so that the robot asks only the questions whose answers actually matter for the plan. When explicit dialogue is unnecessary or impractical, a real-time intent-aware collaboration module instead reads implicit, nonverbal cues, maintaining a probabilistic belief over the human's latent task intent from spatial and directional signals to enable coordination-aware task selection without any communication overhead. We validate the proposed system in both Gazebo simulations and real-world UAV deployments, integrated with a voice dialogue interface and a Vision-Language Model (VLM)-based 3D semantic perception pipeline. Experimental results show that cost-optimal clarification dialogue cuts the interaction cost by 51.9% while maintaining a 100% task success rate, and implicit intent reading reduces the cooperative task execution time by 25.4% compared to the baselines.

2 Citations
0 Influential
3.5 Altmetric
19.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!