2606.06357v1 Jun 04, 2026 cs.SD

F3-토크나이저: 오디오 오토인코더 잠재 변수를 활용하여 이해 및 생성을 가능하게 하는 방법

F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation

Dinghao Zhou
Dinghao Zhou
Citations: 25
h-index: 3
Xingchen Song
Xingchen Song
Citations: 29
h-index: 3
Di Wu
Di Wu
Citations: 32
h-index: 3
Peng Cheng
Peng Cheng
Citations: 11
h-index: 2
Sixian Lv
Sixian Lv
Citations: 5
h-index: 1
Shengfan Shen
Shengfan Shen
Citations: 2
h-index: 1

연속적인 오디오 오토인코더는 파형을 잘 재구성하지만, 종종 이해를 위한 약한 구조의 잠재 변수를 생성하며, 자기 지도 학습 기반 오디오 인코더는 의미론적 정보를 담고 있지만 직접 디코딩하기 어렵습니다. 이러한 불일치는 이해와 생성을 모두 지원해야 하는 단일 오디오 토크나이저 설계에 어려움을 야기합니다. 본 연구에서는 연속적인 오토인코더 잠재 변수를 활용하여 두 가지 구성 요소를 통해 이 문제를 해결합니다: 노이즈 정규화된 오토인코더 병목 구조 및 잠재 변수 측 표현 인코더. 병목 구조는 KL 기반의 변분 학습 대신 채널 정규화와 확률적 교란을 사용하여 재구성 및 자기 회귀 생성을 위한 크기 조절 가능한 연속적인 잠재 변수를 생성합니다. 표현 인코더는 RQ-MTP 및 고정된 LLM(Large Language Model)의 지도 하에, 이미 학습된 오토인코더 잠재 변수에 대해 훈련됩니다. 결과적으로 생성된 토크나이저는 이해를 위한 고차원 표현을 제공하면서 동시에 정규화된 연속적인 잠재 변수를 생성을 위한 타겟으로 유지합니다.

Original Abstract

Continuous audio autoencoders reconstruct waveforms well but often produce latents with weak structure for understanding, while self-supervised audio encoders capture semantics but are not directly decodable. This mismatch complicates a single audio tokenizer that must support both understanding and generation. We adapt continuous autoencoder latents to this setting with two components: a noise-regularized autoencoder bottleneck and a latent-side representation encoder. The bottleneck uses channel normalization and stochastic perturbation instead of KL-based variational training, yielding scale-controlled continuous latents for reconstruction and autoregressive generation. The representation encoder is trained on frozen autoencoder latents with RQ-MTP and frozen-LLM supervision. The resulting tokenizer provides high-dimensional representations for understanding while preserving normalized continuous latents as generation targets

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!