2606.18747v1 Jun 17, 2026 cs.RO

LLM을 활용한 인간 피드백 기반 반복적 강화 학습을 통한 자연스럽고 표현력 있는 로봇 제스처 생성

Generating Natural and Expressive Robot Gestures through Iterative Reinforcement Learning with Human Feedback using LLMs

F. Salim
F. Salim
Citations: 1,590
h-index: 19
Francisco Cruz
Francisco Cruz
Citations: 4
h-index: 1
Chris Lee
Chris Lee
Citations: 9
h-index: 1
Benjamin Tag
Benjamin Tag
Citations: 18
h-index: 3

표현력이 풍부한 제스처는 자연스럽고 효과적인 의사소통에 필수적이며, 언어적 신호만으로는 부족할 때 보완적인 역할을 합니다 (예: 가리키기). 휴머노이드 로봇인 Pepper와 같은 사회용 로봇의 경우, 자연스럽고 표현력 있는 움직임을 생성하는 것은 인간-로봇 상호작용(HRI)을 개선하고 장기적인 수용도를 높이는 데 매우 중요합니다. 그러나 기존 제스처 생성 방식은 전문가가 제작한 애니메이션에 의존하기 때문에 환경 변화에 유연하게 대응하지 못하고 획일화된 동작을 보이며, 이는 역동적이고 다양한 환경에서는 실용적이지 않습니다. 또한, 머신러닝 기반 접근 방식은 자연스러움을 제대로 반영하지 못하는 경우가 많으며, 자유도가 증가할수록 문제가 더욱 심각해집니다. 따라서 표현력 있는 로봇 제스처를 생성하기 위해서는 사회적 규범과 물리적 제약을 준수하면서도 환경에 적응할 수 있는 시스템이 필요합니다. 최근 대규모 언어 모델(LLM)의 발전은 동적인 코드 생성을 가능하게 하여, 자연어를 기반으로 실시간 제스처를 합성하는 새로운 기회를 제공합니다. 본 논문에서는 ChatGPT를 휴머노이드 로봇 Pepper에 통합하여, 대화 내용과 관련된 제스처를 생성합니다. 이 기본 모델은 유연한 제스처 생성이 가능하지만, 결과적으로 생성되는 동작이 종종 뻣뻣하고 부자연스러워 보입니다. 이러한 한계를 극복하기 위해, 사용자 평가를 기반으로 제스처 생성을 개선하는 반복적 강화 학습 with 인간 피드백(RLHF) 시스템을 도입했습니다. 이 시스템은 Pepper가 생성한 제스처에 대한 사용자 연구를 통해 성능을 비교합니다. 실험 결과는 RLHF가 LLM의 동시 음성 제스처 생성 능력을 향상시켜, 더욱 표현력이 풍부하고 관련성이 높으며 자연스러운 움직임을 만들어낸다는 것을 보여줍니다.

Original Abstract

Expressive gestures are essential for natural and effective communication, complementing speech when verbal cues alone are insufficient (e.g., pointing). For social robots such as the humanoid Pepper, producing natural and expressive movements is critical for improving human-robot interaction (HRI) and long-term acceptance. However, generating gestures remains challenging due to reliance on expert-authored animations, resulting in rigid behaviors that are impractical for dynamic and diverse environments. Alternatively, machine learning approaches often struggle to capture perceived naturalness, becoming increasingly challenging with more degrees of freedom. Consequently, producing expressive robot gestures requires a system that can adapt to the environment while adhering to social norms and physical constraints. Recent advances in large language models (LLMs) enable dynamic code generation, offering new opportunities for runtime gesture synthesis from natural language. In this paper, we integrate ChatGPT into the humanoid robot Pepper to generate co-speech gestures aligned with conversational output. While this baseline enables flexible gesture generation, the resulting motions are often perceived as stiff and unnatural. To address this limitation, we introduce an iterative reinforcement learning with human feedback (RLHF) system that finetunes gesture generation based on user evaluations, leveraging an iterative user study to compare Pepper's generated gestures. Our results show that RLHF improved the LLM's co-speech generative capabilities, producing more expressive, relevant and fluid movements.

0 Citations
0 Influential
9.5 Altmetric
47.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!