AG-REPA: 오디오 플로우 매칭에서의 표현 정렬을 위한 인과적 레이어 선택 방법
AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching
REPA (Representation Alignment)는 생성형 플로우 모델 훈련 시, 중간 은닉 상태를 사전 훈련된 교사 모델의 특징과 정렬하여 성능을 향상시키지만, 토큰 기반 오디오 플로우 매칭에서 REPA의 효과는 어떤 감독 레이어를 선택하는지에 크게 의존하며, 이는 일반적으로 깊이를 기준으로 경험적으로 결정됩니다. 본 연구에서는 오디오 플로우 매칭에서의 표현 정렬을 위한 새로운 인과적 레이어 선택 전략인 Attribution-Guided REPA (AG-REPA)를 소개합니다. 첫째, 의미/음향 정보를 가장 잘 저장하는 (교사 공간과의 유사도가 높은) 레이어가 반드시 생성에 기여하는 속도장에 가장 큰 영향을 미치는 레이어가 아님을 발견했으며, 이를 '저장-기여 분리'(SCD)라고 명명했습니다. 이러한 통찰력을 실제 훈련에 적용하기 위해, 각 레이어가 예측된 속도장에 미치는 인과적 기여도를 정량화하는 순방향-전용 게이트 제거(FoG-A) 방법을 제안합니다. 이를 통해 희소 레이어 선택과 정렬을 위한 적응적 가중치를 가능하게 합니다. 다양한 토큰 기반 조건 설정 하에서 통일된 음성 및 일반 오디오 훈련(LibriSpeech + AudioSet)을 수행한 결과, AG-REPA는 일관되게 REPA 기본 모델보다 우수한 성능을 보였습니다. 전반적으로, 본 연구 결과는 정렬이 표현적으로 풍부하지만 기능적으로 비활성적인 레이어보다는 속도장을 주도하는 인과적으로 지배적인 레이어에 적용될 때 가장 효과적이라는 것을 보여줍니다.
REPresentation Alignment (REPA) improves the training of generative flow models by aligning intermediate hidden states with pretrained teacher features, but its effectiveness in token-conditioned audio Flow Matching critically depends on the choice of supervised layers, which is typically made heuristically based on the depth. In this work, we introduce Attribution-Guided REPresentation Alignment (AG-REPA), a novel causal layer selection strategy for representation alignment in audio Flow Matching. Firstly, we find that layers that best store semantic/acoustic information (high teacher-space similarity) are not necessarily the layers that contribute most to the velocity field that drives generation, and we call it Store-Contribute Dissociation (SCD). To turn this insight into an actionable training guidance, we propose a forward-only gate ablation (FoG-A) that quantifies each layer's causal contribution via the induced change in the predicted velocity field, enabling sparse layer selection and adaptive weighting for alignment. Across unified speech and general-audio training (LibriSpeech + AudioSet) under different token-conditioning topologies, AG-REPA consistently outperforms REPA baselines. Overall, our results show that alignment is most effective when applied to the causally dominant layers that drive the velocity field, rather than to layers that are representationally rich but functionally passive.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.