2608.07631v1 Aug 07, 2026 cs.SD

PACE: LLM 기반 양방향 음성 대화 시스템을 위한 재생에 정렬된 컨텍스트 엔진

PACE: A Playback-Aligned Context Engine for LLM-Based Full-Duplex Voice Dialogue

Shibo Wang
Shibo Wang
Citations: 8,020
h-index: 10
Junfeng Ma
Junfeng Ma
Citations: 3
h-index: 1
Zicheng Zhang
Zicheng Zhang
Citations: 0
h-index: 0
Libo Wang
Libo Wang
Citations: 4
h-index: 1

LLM(대규모 언어 모델) 기반의 양방향 음성 서비스는 사용자가 응답하는 동안에도 계속해서 말할 수 있도록 지원합니다. 서버가 클라이언트보다 빠르게 응답을 생성하고 대화 상태를 업데이트하기 때문에, 사용자가 실제로 듣지 못한 내용에 기반하여 후속 발언이 해석될 수 있습니다. 우리는 이러한 문제를 '생성된 컨텍스트의 부정확한 참조(Generative Context Mis-anchoring, GCM)'라고 부릅니다. GCM 문제를 해결하기 위해, 저희는 모델과 상호작용하는 컨텍스트를 클라이언트 재생 경계에 연결하는, 즉 사용자가 실제로 들을 수 있는 내용에 대한 시스템적으로 관찰 가능한 프록시인 'PACE'라는 독립적인 미들웨어 레이어를 제안합니다. PACE는 인터럽트 발생 후, 재생되지 않은 어시스턴트의 내용을 제외하여 컨텍스트를 복구하고, 다양한 음성 런타임 환경에서도 낮은 지연 시간으로 응답을 생성할 수 있도록 합니다. 저희는 블랙박스 음성 모델을 사용하여 브라우저 기반 실시간 음성 어시스턴트에서 PACE의 오디오 전용 경로를 전체적으로 구현했으며, 모델 서비스 자체를 수정하지 않았습니다. 또한, 108개의 재생 관련 참조 고정 사례로 구성된 새로운 제어 벤치마크 데이터셋인 'GCM-Bench'를 구축했습니다. GCM-Bench에서 PACE는 참조 고정 정확도를 25.0%에서 96.3%까지 향상시켰습니다. 또한, 200개의 Full-Duplex-Bench v1 인터럽트 샘플에 대해 인터럽트 응답 품질을 유지했습니다. 이러한 결과는 모델이 사용하는 컨텍스트를 실제 재생 내용에 연결하는 것이 양방향 음성 대화의 일관성을 유지하는 실용적인 방법임을 보여줍니다.

Original Abstract

LLM-based full-duplex voice services allow users to speak while the assistant is responding. Because servers can generate output and advance dialogue state faster than clients can play it, subsequent user speech may be interpreted based on content the user never heard. We call this failure Generative Context Mis-anchoring (GCM). To address GCM issues, we present PACE, a provider-independent middleware layer that anchors model-facing context to the client playback boundary, a system-observable proxy for what the user could have heard. After an interruption, PACE repairs this context to exclude assistant content that never reached playback, while preserving low-latency generation across heterogeneous voice runtimes. We implement PACE's audio-only projection path end to end in a browser-based realtime voice assistant using a black-box speech model, without modifying the model service. We also construct GCM-Bench, a new controlled benchmark dataset of 108 playback-relative referent-anchoring cases. On GCM-Bench, PACE raises Referent Anchoring Accuracy from 25.0% to 96.3% over a cancellation-only baseline. On 200 Full-Duplex-Bench v1 interruption samples, it preserves interruption response quality. These results show that grounding model-facing context in actual playback is a practical way to maintain consistency in full-duplex voice dialogue.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!