UniSAE: 화자, 감정 및 저수준 콘텐츠에 대한 통합 음성 속성 편집 - 이산 음운 후행문 모델링을 통한 접근
UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling
음성 편집은 발화의 특정 부분을 수정하면서 나머지 음성을 보존하는 것을 목표로 합니다. 기존 연구는 주로 단어 수준의 내용 수정에 초점을 맞추며, 화자, 감정 및 내용 편집을 별개의 작업으로 취급하여 편집의 세밀함과 유연성을 제한합니다. 본 논문에서는 하위 음소부터 단어 수준까지 다양한 범위를 아우르는 통합된 화자, 감정 및 내용 편집 기능을 제공하는 UniSAE 프레임워크를 제안합니다. UniSAE는 음성 콘텐츠를 음소 식별, 발음 변형 및 지속 시간을 인코딩하는 이산 토큰으로 분해하는 이산 음운 후행문(DPPG) 표현 방식을 도입하여 직접적인 음소 및 하위 음소 수준의 편집을 가능하게 합니다. 보다 상위 수준의 수정 작업을 위해, 자기 회귀 콘텐츠 트랜스포머는 단어 수준의 내용 편집을 위한 편집된 DPPG 시퀀스를 예측합니다. 편집된 시퀀스는 분리된 화자 및 감정 표현에 조건부로 작동하는 확산 기반 음향 디코더를 통해 음성으로 렌더링됩니다. 실험 결과는 제안된 통합 프레임워크가 정밀한 화자 및 감정 제어, 다양한 수준의 내용 편집, 그리고 단일 프레임워크 내에서 세 가지 속성의 공동 수정 기능을 지원함을 보여줍니다.
Speech editing aims to modify specific portions of an utterance while preserving the remaining speech. Existing approaches primarily focus on word-level content modification and typically treat content, speaker, and emotion editing as separate tasks, limiting both editing granularity and flexibility. We propose UniSAE, a unified speech attribute editing framework which supports composable speaker, emotion and content editing from sub-phoneme to word level within a single architecture. UniSAE introduces a Discrete Phonetic PosteriorGram (DPPG) representation that factorizes speech content into discrete tokens encoding phoneme identity, pronunciation variants, and duration, enabling direct phoneme- and sub-phoneme-level editing. For higher-level modifications, an autoregressive content transformer predicts edited DPPG sequences for word-level content editing. The edited sequences are rendered into speech by a diffusion-based acoustic decoder, conditioned on disentangled speaker and emotion representations. Experimental results demonstrate that the proposed unified framework supports precise speaker and emotion control, content editing at multiple granularities, and joint modification of all three attributes within a single framework.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.