HoliTok: 연속적인 통합 토큰화 모델 – 음성 생성 및 이해를 위한 강력한 이중 기능
HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding
통합된 음성 기반 모델은 언어 모델이 학습할 수 있을 뿐만 아니라 고품질 웨이브폼으로 디코딩될 수 있는 통합적인 토큰화 공간을 요구합니다. 그러나 기존의 음성 토크나이저는 종종 이러한 요구 사항을 동시에 만족시키지 못하여, 아키텍처 복잡성이 증가하고 훈련 설계가 더욱 복잡해지는 경향이 있습니다. 본 논문에서는 통합 생성-이해 모델링을 위해 설계된 연속적인 통합 음성 토큰화 모델인 HoliTok을 제안합니다. HoliTok은 48kHz 음성을 128차원 잠재 벡터 시퀀스로 압축하여, 25Hz 간격으로 표현합니다. 이 모델은 신호 수준의 충실도를 유지하고, 의미 정보를 통합하며, 강력한 잠재 학습 가능성을 갖도록 점진적인 전략을 통해 훈련됩니다. 이러한 토큰화를 기반으로, 음성 합성 및 인식에 사용되는 통합 AR+DiT 모델을 구축했습니다. 동일한 잠재 벡터 시퀀스가 생성 특화 작업과 통합된 생성-이해 작업 모두를 지원합니다. 실험 결과, HoliTok은 경쟁력 있는 재구성 정확도를 달성하고, 고품질의 제어 가능한 합성을 위한 생성 학습 능력을 향상시킵니다. 또한, 평가된 표현 방식 중 HoliTok만이 추가적인 최적화 트릭 없이 통합된 생성-이해 아키텍처에서 안정적으로 작동하는 것으로 나타났습니다. 이러한 결과는 HoliTok이 효과적인 음성 토크나이저이자 통합된 음성 언어 모델링을 위한 기초적인 표현 인터페이스로 활용될 수 있음을 시사합니다. 코드: https://github.com/bovod-sjtu/HoliTok.
Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-quality waveforms. Existing speech tokenizers, however, often fail to satisfy these requirements simultaneously, leading to increased architectural complexity and more involved training designs. We propose HoliTok, a continuous Holistic speech Tokenization model designed for unified generation-understanding modeling. HoliTok encodes 48~kHz speech into a compact 25~Hz sequence of 128-dimensional latents. It is trained with a progressive strategy that jointly preserves signal-level fidelity, incorporates semantic information, and maintains strong latent learnability. Based on this tokenization, we build a unified AR+DiT model for speech synthesis and recognition, where the same latent sequence supports both generation-specific and unified generation-understanding tasks. Experiments show that HoliTok achieves competitive reconstruction fidelity, improves generative learnability for high-quality and controllable synthesis, and, among the evaluated representations, is the only one that operates robustly in our unified generation-understanding architecture without additional optimization tricks. These results suggest that HoliTok serves as an effective speech tokenizer and a foundational representation interface for unified spoken language modeling. The code is available at: https://github.com/bovod-sjtu/HoliTok.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.