2606.31128v1 Jun 30, 2026 cs.SD

UniSAE: 화자, 감정 및 저수준 콘텐츠에 대한 통합 음성 속성 편집 - 이산 음운 후행문 모델링을 통한 접근

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling

Wei Xue
Wei Xue
Citations: 22
h-index: 2
Yi-Ting Guo
Yi-Ting Guo
Citations: 2,143
h-index: 26
Chuanbo Zhu
Chuanbo Zhu
Citations: 189
h-index: 4
Wuyou Zhou
Wuyou Zhou
Citations: 0
h-index: 0
Rongxiu Zhong
Rongxiu Zhong
Citations: 57
h-index: 3
Shilei Zhang
Shilei Zhang
Citations: 70
h-index: 3
Kun Qian
Kun Qian
Citations: 61
h-index: 3

음성 편집은 발화의 특정 부분을 수정하면서 나머지 음성을 보존하는 것을 목표로 합니다. 기존 연구는 주로 단어 수준의 내용 수정에 초점을 맞추며, 화자, 감정 및 내용 편집을 별개의 작업으로 취급하여 편집의 세밀함과 유연성을 제한합니다. 본 논문에서는 하위 음소부터 단어 수준까지 다양한 범위를 아우르는 통합된 화자, 감정 및 내용 편집 기능을 제공하는 UniSAE 프레임워크를 제안합니다. UniSAE는 음성 콘텐츠를 음소 식별, 발음 변형 및 지속 시간을 인코딩하는 이산 토큰으로 분해하는 이산 음운 후행문(DPPG) 표현 방식을 도입하여 직접적인 음소 및 하위 음소 수준의 편집을 가능하게 합니다. 보다 상위 수준의 수정 작업을 위해, 자기 회귀 콘텐츠 트랜스포머는 단어 수준의 내용 편집을 위한 편집된 DPPG 시퀀스를 예측합니다. 편집된 시퀀스는 분리된 화자 및 감정 표현에 조건부로 작동하는 확산 기반 음향 디코더를 통해 음성으로 렌더링됩니다. 실험 결과는 제안된 통합 프레임워크가 정밀한 화자 및 감정 제어, 다양한 수준의 내용 편집, 그리고 단일 프레임워크 내에서 세 가지 속성의 공동 수정 기능을 지원함을 보여줍니다.

Original Abstract

Speech editing aims to modify specific portions of an utterance while preserving the remaining speech. Existing approaches primarily focus on word-level content modification and typically treat content, speaker, and emotion editing as separate tasks, limiting both editing granularity and flexibility. We propose UniSAE, a unified speech attribute editing framework which supports composable speaker, emotion and content editing from sub-phoneme to word level within a single architecture. UniSAE introduces a Discrete Phonetic PosteriorGram (DPPG) representation that factorizes speech content into discrete tokens encoding phoneme identity, pronunciation variants, and duration, enabling direct phoneme- and sub-phoneme-level editing. For higher-level modifications, an autoregressive content transformer predicts edited DPPG sequences for word-level content editing. The edited sequences are rendered into speech by a diffusion-based acoustic decoder, conditioned on disentangled speaker and emotion representations. Experimental results demonstrate that the proposed unified framework supports precise speaker and emotion control, content editing at multiple granularities, and joint modification of all three attributes within a single framework.

0 Citations
0 Influential
13 Altmetric
65.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!