빠르게 판단하고 신중하게 결정: 확산 모델 기반의 다중 참여자 발화 교대 시스템
Fast When, Careful Who: Dual-Process Multiparty Turn-Taking with Diffusion Augmentation
안정적인 발화 교대는 음성 대화 시스템에서 필수적입니다. 그러나 대부분의 기존 방법은 두 명의 화자 간 상호 작용을 위해 설계되었으며, 겹침과 급격한 화자 변경이 포함된 실제 다중 참여 오디오 환경에서는 어려움을 겪습니다. 본 연구는 VoxConverse 데이터셋을 사용하여 다중 참여 발화 교대를 연구하고, 발화 경계 시점을 판단하는 것과 실제로 발언권이 이전되는 것을 분리하는 두 단계의 오디오 기반 파이프라인을 제안합니다. 빠른 트리거(trigger) 모듈은 오디오를 스캔하여 발화 종료 후보 시간을 제안하며, 가벼운 검증기(verifier)는 해당 시점에만 실행되어 extsc{Hold} 또는 extsc{Shift} 여부를 결정하고 다음 화자 예측을 지원합니다. 본 연구에서는 전체 다중 참여 환경과 비교 가능성을 위해 설계된 통제된 이원화(dyadic) 상위 2개 화자 기반 실험 결과를 보고합니다. 또한, 데이터 증강 전략으로 레이블 보존 기반의 확산 모델을 활용한 배경 오디오 혼합 방법을 조사했습니다. 결과는 기준 모델 대비 개선된 발화 전환 감지 성능을 보여주었으며, 확산 모델 증강 기법을 통해 추가적인 성능 향상을 확인할 수 있었습니다.
Reliable turn-taking is essential for spoken dialogue systems. However, most existing methods are designed for two-speaker interaction and struggle with realistic multiparty audio containing overlap and rapid speaker changes. We study multiparty turn-taking on the VoxConverse dataset and propose an audio-only two-stage pipeline that separates when to trigger a turn boundary from whether the floor is actually transferring. A fast trigger scans the audio and proposes candidate end-of-turn times, while a lightweight verifier runs only at those times to decide \textsc{Hold} or \textsc{Shift} and support next-speaker prediction. We report results in the full multiparty setting and a controlled dyadic top-2 projection for comparability. We also investigate diffusion-based, label-preserving background-audio mixing as a data augmentation strategy. Results show improved shift detection over a baseline, with further improvements from diffusion augmentation.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.