2607.06540v4 Jul 07, 2026 cs.CL

계층적 음성-의미 모델링: 모달 분리 및 의미 일관성을 통한 풀-듀플렉스 음성 언어 모델

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs

Baotian Hu
Baotian Hu
Citations: 388
h-index: 10
Xuanyu Zhang
Xuanyu Zhang
Citations: 152
h-index: 4
Yancheng He
Yancheng He
Citations: 834
h-index: 11
Yunxin Li
Yunxin Li
Harbin Institute of Technology, Shenzhen
Citations: 1,113
h-index: 13
Min Zhang
Min Zhang
Citations: 432
h-index: 8
Zhenyu Liu
Zhenyu Liu
Citations: 387
h-index: 10
Qixun Teng
Qixun Teng
Citations: 6
h-index: 2
Shenyuan Jiang
Shenyuan Jiang
Citations: 241
h-index: 4
Haolan Chen
Haolan Chen
Citations: 26
h-index: 3
Mingjun Zhao
Mingjun Zhao
Citations: 255
h-index: 4
Fanbo Meng
Fanbo Meng
Citations: 31
h-index: 4
Yu Xu
Yu Xu
Citations: 10
h-index: 3
Haizhou Li
Haizhou Li
Citations: 0
h-index: 0

원활하고 고성능의 지능형 풀-듀플렉스 음성 언어 모델(SLM)을 개발하는 것은 음성 처리 및 자연어 처리 분야에서 중요한 과제이자 오랜 목표입니다. 상당한 발전이 있었음에도 불구하고, 최근 연구는 심각한 모달 간섭으로 인해 근본적인 제약을 받으며, 이는 지식 저하를 초래하고 의미적 완전성을 훼손하여 풀-듀플렉스 SLM을 부자연스럽고 비지능적으로 보이게 만듭니다. 본 논문에서는 모델 최적화 과정을 면밀하게 분석하여 이러한 성능 저하의 근본 원인을 밝히고, 음성 및 의미 모델링 간의 깊은 파라미터 공간 공유로 인해 발생하는 모달 간섭이 상반되는 기울기 충돌에서 비롯된다는 것을 보여줍니다. 이 중요한 통찰력을 바탕으로, 우리는 모달 간섭을 완화하도록 설계된 풀-듀플렉스 프레임워크인 Lychee-FD를 소개합니다. 특히, 우리는 심층 레이어에서 충돌하는 모달을 분리하면서 전용 의미 정렬 채널을 통해 모달 간의 일관성을 유지하는 계층적 파라미터 분리 전략을 제안합니다. 여러 풀-듀플렉스 벤치마크에 대한 광범위한 실험 결과, 우리의 방법이 최첨단 기술을 크게 발전시키고 음성 지능(Spoken QA에서 +7.4% 향상) 및 풀-듀플렉스 상호 작용의 유창성(FullDuplexBench 1.5에서 +28.5% 향상) 모두에서 상당한 개선을 가져왔으며, 추론 효율성을 저해하지 않는다는 것을 보여줍니다. 현재까지 알려진 바로는, 본 연구는 다음과 같은 두 가지 주요 진전을 이루었습니다. 첫째, 풀-듀플렉스 SLM에서 발생하는 모달 간섭의 근본 원인을 밝히고 설명했으며, 둘째, 원활하고 고성능의 지능형 풀-듀플렉스 SLM을 위한 우아한 계층적 모델과 실용적인 솔루션을 설계했습니다.

Original Abstract

Developing seamless, high-performance, native intelligent full-duplex Spoken Language Models (SLMs) remains a critical challenge and long-standing goal for the speech and NLP community. Despite notable progress, recent endeavors are fundamentally constrained by severe modality interference, which causes substantial knowledge degradation and compromises semantic integrity -- ultimately making full-duplex SLMs feel unnatural and unintelligent. In this paper, through an exhaustive fine-grained analysis of model optimization dynamics, we uncover the root cause of such performance degradation, revealing that modality interference arises from inherent gradient conflicts between acoustic and semantic modeling when the two modalities are forced to share a deep parameter space. Guided by this key insight, we introduce Lychee-FD, a native end-to-end full-duplex framework designed to mitigate modality interference. Importantly, we propose a hierarchical parameter separation strategy that decouples conflicting modalities in deep layers while preserving cross-modality coherence via a dedicated semantic alignment channel. Extensive experiments on multiple full-duplex benchmarks demonstrate that our method significantly advances the state of the art, yielding substantial improvements in both speech intelligence (+7.4% on Spoken QA) and full-duplex interaction fluidity (+28.5% on FullDuplexBench 1.5) without compromising inference efficiency. To the best of our knowledge, this work is the first to achieve two key advances: 1) uncovering and elucidating the root cause of modality interference in full-duplex SLMs, and 2) designing an elegant hierarchical model together with a practical solution for seamless, high-performance, native intelligent full-duplex SLMs.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!