ProPS: 자연어 기반 음성 특징 분포 생성을 위한 프롬프트 기반 프로필 합성
ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions
음성 특징(x-벡터)은 화자 식별 및 화자와 관련된 속성을 나타내는 데 널리 사용되지만, 기존의 특징 추출기는 일반적으로 기술적인 역할을 수행합니다. 즉, 관찰된 음성 세그먼트를 x-벡터로 매핑한 후 이를 추가적인 작업에 활용합니다. 본 논문에서는 자연어 프롬프트(예: “인도 악센트가 있는 30대 남성 화자”)에 따라 화자 특징 분포를 생성하는 프레임워크인 ProPS (Prompted Profile Synthesis)를 소개합니다. ProPS는 사람이 작성한 프로필 설명을 문장 임베딩으로 변환하고, 대규모 데이터 세트에 대해 학습된 혼합 밀도 네트워크(Mixture Density Network)를 사용하여 x-벡터 공간에서 가우스 혼합 모델을 예측합니다. 이 모델은 실제 화자 특징이 요청된 프로필과 일치할 가능성을 최대화하도록 학습되며, 생성된 분포는 보류된 x-벡터에 대한 음의 로그 우도를 기준으로 평가되고, 샘플링된 합성 x-벡터에 대한 속성 분류 정확도를 통해 검증됩니다. 실험 결과, ProPS는 프로필 조건화된 분포를 생성하며, 나이, 성별, 억양 및 발음 특징과 같은 요청된 화자 속성을 유지하는 x-벡터를 생성합니다. 이러한 설계는 Text-To-Speech (TTS) 또는 음성 변환(Voice Conversion)과 같은 음성 생성 시스템에서 제어 가능한 화자 프로필 합성을 가능하게 하며, 동시에 생성된 분포를 관찰된 화자 특징 구조에 고정시키는 역할을 합니다.
Speaker embeddings, or x-vectors, are widely used to represent speaker identity and speaker-related attributes, but existing embedding extractors are typically descriptive rather than generative: they map an observed speech segment to an x-vector, which is then used for downstream applications. We introduce ProPS, Prompted Profile Synthesis, a framework for generating distributions of speaker embeddings conditioned on natural language prompts such as "a thirties male speaker with an Indian accent". ProPS converts human-written profile descriptions into sentence embeddings and uses a mixture density network trained on a large-scale dataset to predict a Gaussian mixture model in the x-vector space. The model is trained by maximizing the likelihood that real speaker embeddings match the requested profile, and its generated distributions are evaluated by negative log-likelihood on held-out x-vectors and by attribute classification accuracies on sampled synthetic x-vectors. Experiments show that ProPS produces profile-conditioned distributions and generates x-vectors that preserve requested speaker attributes such as age, gender, accent, and prosodic characteristics. This design enables controllable speaker-profile synthesis for speech generation systems like Text-To-Speech (TTS) or Voice Conversion (VC) while anchoring generated distributions in observed speaker-embedding structure.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.