실시간 다자간 음성 에이전트를 위한 적응형 발화 교대 방식
Adaptive Turn-Taking for Real-time Multi-Party Voice Agents
다자간 음성 대화에서 발화 교대 방식은 특히 동적인 발화 권리 경쟁과 다양한 사용자 기대치 하에서 음성 기반 에이전트에게 여전히 중요한 과제입니다. 본 논문에서는 다자간 환경에서 명시적으로 할당된 역할에 따라 발화 교대 행동을 조절하는 롤플레잉 음성 에이전트인 ModeratorLM을 제안합니다. 이 시스템은 청크 단위 스트리밍 방식으로 작동하는 음성 기반 대규모 언어 모델을 기반으로 구축되었습니다. 또한, 대화 맥락과 할당된 역할에 대한 연쇄적 사고(chain-of-thought reasoning)를 통합한 강화된 변형을 소개합니다. 우리는 다양한 어시스턴트 역할을 가진 다자간 음성 대화를 포함하는 대규모 합성 데이터셋인 RolePlayConv을 구축했습니다. 실제 회의 데이터 및 RolePlayConv에서의 실험 결과, 제안하는 방식은 역할 기반이 아닌 기준 모델에 비해 발화 교대 정확도를 40% 이상, 재현율을 70% 이상 향상시키는 동시에 오탐으로 인한 불필요한 중단을 크게 줄이는 것을 확인했습니다.
Turn-taking in multi-party spoken conversations remains a fundamental challenge for voice-based agents, particularly under dynamic floor competition and varying user expectations. We propose ModeratorLM, a role-playing voice agent that conditions turn-taking behavior on an explicitly assigned role in multi-party settings. The system is built on a speech large language model operating in chunk-wise streaming manner. We further introduce a reasoning-augmented variant that incorporates chain-of-thought reasoning over conversational context and the assigned role. We construct RolePlayConv, a large-scale synthetic dataset of spoken multi-party conversations with diverse assistant roles. Experiments on real-world meeting data and RolePlayConv show improved turn-taking precision by over 40% and recall by more than 70%, while substantially reducing false-positive interruptions compared to non-role-conditioned baselines.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.