2607.27756v1 Jul 30, 2026 cs.SD

칵테일 토커(Cocktail-Talker): 노이즈가 많은 사회적 환경에서의 다중 화자 대화 모델링 - 턴 액션 GRPO 활용

Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO

N. Mesgarani
N. Mesgarani
Citations: 13,484
h-index: 47
Xilin Jiang
Xilin Jiang
Citations: 315
h-index: 8
Riki Shimizu
Riki Shimizu
Citations: 10
h-index: 2
Sukru Samet Dindar
Sukru Samet Dindar
Citations: 63
h-index: 3
Junkai Wu
Junkai Wu
University of Washington
Citations: 55
h-index: 4
Zhongweiyang Xu
Zhongweiyang Xu
Citations: 29
h-index: 3

음성 대화 시스템은 일반적으로 깨끗하고 양방향적인 상호작용을 위해 설계되며, 하나의 사용자(또는 사용자와 어시스턴트)가 번갈아 가며 말하는 방식으로 구성됩니다. 그러나 실제 사회적 대화는 종종 더 복잡하며, 여러 화자가 관련 없는 발언과 배경 소음 속에서 동시에 참여할 수 있습니다. 각 발언은 어시스턴트를 대상으로 할 수도 있고, 다른 화자에게 전달될 수도 있으며, 또는 완전히 무관한 내용일 수도 있습니다. 이러한 환경에서 어시스턴트는 무엇을 말해야 하는지뿐만 아니라, вообще 말을 해야 하는지 여부를 결정해야 합니다. 본 논문에서는 노이즈가 많은 사회적 환경에서의 다중 화자 음성 대화 모델링을 위한 언어 모델 프레임워크인 '칵테일 토커(Cocktail-Talker)'를 소개합니다. 우리는 어시스턴트의 행동을 <|respond|> (응답), <|listen|> (경청), 및 <|ignore|> (무시)라는 세 가지 액션 토큰으로 모델링하며, 이 토큰들은 응답 또는 침묵 앞에 위치합니다. 칵테일 토커는 지도 학습과 강화 학습을 통해 적절한 액션 토큰을 생성하도록 훈련되며, <|respond|> 모드에서만 음성 응답을 생성합니다. 훈련 데이터를 준비하기 위해, 우리는 다양한 사회적 상황에서 화자 역할을 시뮬레이션하는 현실적인 다중 화자 대화를 생성하는 LLM 기반 데이터 파이프라인인 '칵테일-다이얼로그젠(Cocktail-DialogGen)'을 개발했습니다. 이러한 구성 요소들을 통해, 복잡한 사회적 환경에서 보다 자연스럽고 선택적으로 상호 작용하는 음성 대화 시스템으로 나아갈 수 있습니다.

Original Abstract

Spoken dialog systems are typically designed for clean, dyadic interactions in which a single user and an assistant take turns speaking. Real-world social conversations, however, are often more ambiguous: multiple speakers may participate in the same conversation amid irrelevant speech and background noise. Each utterance may be directed to the assistant, addressed to another speaker, or completely irrelevant. In such settings, the assistant must decide not only what to say, but also whether to speak at all. In this paper, we introduce Cocktail-Talker, a speech LLM framework for multi-speaker spoken dialog modeling in noisy social environments. We model the assistant's behavior with three action tokens: <|respond|>, <|listen|>, and <|ignore|>, placed before a response or silence. Cocktail-Talker is trained via supervised finetuning and reinforcement learning to generate the appropriate action token and, only in <|respond|> mode, a speech response. To prepare the training data, we develop Cocktail-DialogGen, an LLM-based data pipeline that simulates realistic multi-speaker dialogs with speaker roles across diverse social settings. Together, these components take a step toward spoken dialog systems that interact more naturally and selectively in complex social environments.

0 Citations
0 Influential
23.5 Altmetric
117.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!