지시문이 음성에 미치는 영향: 스타일 정보가 포함된 텍스트 음성 변환 모델의 교차 어텐션 분석
How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech
스타일 정보를 활용한 텍스트 음성 변환 시스템은 자연어 설명을 통해 음색 특성을 제어하지만, 개별 단어가 음향 출력에 미치는 영향은 명확하지 않습니다. 이러한 이해는 오류 원인을 진단하고 표현력이 풍부한 TTS 모델의 제어 가능성을 향상시키는 데 매우 중요합니다. 본 연구에서는 음성 확산 모델에 대한 교차 어텐션 분석 방법을 제안하며, DAAM 프레임워크를 음성 분야에 처음으로 적용하여 CapSpeech-TTS 모델에 적용했습니다. 우리 방법은 25개의 레이어와 24개의 ODE 단계를 통해 토큰별 히트맵을 추출합니다. 총 3,600개의 (스타일 설명, 텍스트 전사) 조합(각 스타일 설명이 30개의 텍스트 전사를 생성하는 조건)을 분석하여, 스타일 설명의 토큰이 음파형에 미치는 영향을 파악했습니다. 연구 결과는 다음과 같습니다: (1) 스타일 토큰은 내용/기능 토큰보다 시간적 변동성이 낮아 전체적인 제어를 확인시켜줍니다; (2) 스타일 어텐션은 기본 주파수(F0) 및 에너지와 상관관계가 있습니다; (3) 스타일 조건화는 초기 단계 및 깊은 레이어에서 가장 두드러집니다; (4) 어텐션 엔트로피는 17번째 레이어에서 최소값을 가지며, 이는 스타일 중요성이 최고조에 달하는 시점과 일치하며, 네트워크의 선택성이 가장 중요한 단계에서 극대화됨을 나타냅니다. 본 연구는 자연어가 음성 확산 모델의 교차 어텐션에 미치는 영향을 분석한 최초의 사례입니다.
Style-captioned text-to-speech systems use natural language to control voice characteristics, but how individual words influence acoustic output remains unclear. Understanding this is critical for diagnosing failure modes and improving controllability in expressive TTS. We propose cross-attention attribution for speech diffusion models, adapting the DAAM framework to the speech domain for the first time, and apply it to CapSpeech-TTS. Our method extracts per-token heatmaps across 25 layers and 24 ODE steps. We analyze 3,600 (style caption, text transcript) combinations comprising 120 style captions conditioning the generation of 30 text transcripts each, revealing how caption tokens shape waveforms. Results show: (1) style tokens have lower temporal variance than content/function tokens, confirming global conditioning; (2) style attention correlates with F0 and energy; (3) style conditioning peaks in early steps and deep layers; (4) attention entropy reaches its minimum at layer 17, co-occurring with the style importance peak, indicating maximal network selectivity at the most style-critical stage. This is the first study of how natural language influences cross-attention in speech diffusion models
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.