2606.20532v1 Jun 18, 2026 cs.AI

지시문이 음성에 미치는 영향: 스타일 정보가 포함된 텍스트 음성 변환 모델의 교차 어텐션 분석

How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech

Nityanand Mathur
Nityanand Mathur
IIIT Guwahati, Bosch Research
Citations: 16
h-index: 2
Akshat Mandloi
Akshat Mandloi
Citations: 1
h-index: 1
H. Sayed
H. Sayed
Citations: 4
h-index: 2
Wasim Madha
Wasim Madha
Citations: 0
h-index: 0
A. Singh
A. Singh
Citations: 0
h-index: 0
Sameer Khurana
Sameer Khurana
Citations: 113
h-index: 6
Sudarshan Kamath
Sudarshan Kamath
Citations: 0
h-index: 0

스타일 정보를 활용한 텍스트 음성 변환 시스템은 자연어 설명을 통해 음색 특성을 제어하지만, 개별 단어가 음향 출력에 미치는 영향은 명확하지 않습니다. 이러한 이해는 오류 원인을 진단하고 표현력이 풍부한 TTS 모델의 제어 가능성을 향상시키는 데 매우 중요합니다. 본 연구에서는 음성 확산 모델에 대한 교차 어텐션 분석 방법을 제안하며, DAAM 프레임워크를 음성 분야에 처음으로 적용하여 CapSpeech-TTS 모델에 적용했습니다. 우리 방법은 25개의 레이어와 24개의 ODE 단계를 통해 토큰별 히트맵을 추출합니다. 총 3,600개의 (스타일 설명, 텍스트 전사) 조합(각 스타일 설명이 30개의 텍스트 전사를 생성하는 조건)을 분석하여, 스타일 설명의 토큰이 음파형에 미치는 영향을 파악했습니다. 연구 결과는 다음과 같습니다: (1) 스타일 토큰은 내용/기능 토큰보다 시간적 변동성이 낮아 전체적인 제어를 확인시켜줍니다; (2) 스타일 어텐션은 기본 주파수(F0) 및 에너지와 상관관계가 있습니다; (3) 스타일 조건화는 초기 단계 및 깊은 레이어에서 가장 두드러집니다; (4) 어텐션 엔트로피는 17번째 레이어에서 최소값을 가지며, 이는 스타일 중요성이 최고조에 달하는 시점과 일치하며, 네트워크의 선택성이 가장 중요한 단계에서 극대화됨을 나타냅니다. 본 연구는 자연어가 음성 확산 모델의 교차 어텐션에 미치는 영향을 분석한 최초의 사례입니다.

Original Abstract

Style-captioned text-to-speech systems use natural language to control voice characteristics, but how individual words influence acoustic output remains unclear. Understanding this is critical for diagnosing failure modes and improving controllability in expressive TTS. We propose cross-attention attribution for speech diffusion models, adapting the DAAM framework to the speech domain for the first time, and apply it to CapSpeech-TTS. Our method extracts per-token heatmaps across 25 layers and 24 ODE steps. We analyze 3,600 (style caption, text transcript) combinations comprising 120 style captions conditioning the generation of 30 text transcripts each, revealing how caption tokens shape waveforms. Results show: (1) style tokens have lower temporal variance than content/function tokens, confirming global conditioning; (2) style attention correlates with F0 and energy; (3) style conditioning peaks in early steps and deep layers; (4) attention entropy reaches its minimum at layer 17, co-occurring with the style importance peak, indicating maximal network selectivity at the most style-critical stage. This is the first study of how natural language influences cross-attention in speech diffusion models

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!