DuplexCascade: VAD-Free 캐스케이드 ASR-LLM-TTS 파이프라인과 마이크로 턴 최적화를 통한 풀-듀플렉스 음성-음성 대화 시스템
DuplexCascade: Full-Duplex Speech-to-Speech Dialogue with VAD-Free Cascaded ASR-LLM-TTS Pipeline and Micro-Turn Optimization
캐스케이드 ASR-LLM-TTS 모듈을 사용하는 음성 대화 시스템은 강력한 LLM(Large Language Model)의 지능을 유지하지만, VAD(Voice Activity Detection) 분할은 종종 반-듀플렉스 방식으로 작동하게 하고, 시스템의 제어 방식을 불안정하게 만듭니다. 반면, VAD를 사용하지 않는 엔드-투-엔드 모델은 풀-듀플렉스 상호 작용을 지원하지만, 대화의 지능을 유지하기 어렵습니다. 본 논문에서는 VAD-Free 캐스케이드 스트리밍 파이프라인인 DuplexCascade를 제안합니다. 핵심 아이디어는 기존의 발화 단위의 긴 대화 흐름을 덩어리 단위의 마이크로 턴 상호 작용으로 변환하여 빠른 양방향 교환을 가능하게 하면서 동시에 강력한 텍스트 LLM의 장점을 유지하는 것입니다. 신뢰할 수 있는 턴-테이킹(Turn-taking) 및 응답 타이밍 조정을 위해, LLM의 동작을 스트리밍 제약 조건 하에서 제어하는 데 사용되는 특수한 제어 토큰 세트를 도입했습니다. Full-DuplexBench 및 VoiceBench 데이터셋에서 DuplexCascade는 최첨단 풀-듀플렉스 턴-테이킹 성능과 강력한 대화 지능을 보여주며, 오픈 소스 음성-음성 대화 시스템 중에서 우수한 성능을 달성했습니다.
Spoken dialog systems with cascaded ASR-LLM-TTS modules retain strong LLM intelligence, but VAD segmentation often forces half-duplex turns and brittle control. On the other hand, VAD-free end-to-end model support full-duplex interaction but is hard to maintain conversational intelligence. In this paper, we present DuplexCascade, a VAD-free cascaded streaming pipeline for full-duplex speech-to-speech dialogue. Our key idea is to convert conventional utterance-wise long turns into chunk-wise micro-turn interactions, enabling rapid bidirectional exchange while preserving the strengths of a capable text LLM. To reliably coordinate turn-taking and response timing, we introduce a set of conversational special control tokens that steer the LLM's behavior under streaming constraints. On Full-DuplexBench and VoiceBench, DuplexCascade delivers state-of-the-art full-duplex turn-taking and strong conversational intelligence among open-source speech-to-speech dialogue systems.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.