2605.29948v1 May 28, 2026 cs.SD

HoliTok: 연속적인 통합 토큰화 모델 – 음성 생성 및 이해를 위한 강력한 이중 기능

HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

Hankun Wang
Hankun Wang
Citations: 264
h-index: 7
Kai Yu
Kai Yu
Citations: 176
h-index: 6
Yiwei Guo
Yiwei Guo
Citations: 755
h-index: 14
Colin Zhang
Colin Zhang
Citations: 85
h-index: 6
Shiyue Lian
Shiyue Lian
Citations: 1
h-index: 1
Yu Xi
Yu Xi
Citations: 166
h-index: 6
Da Zheng
Da Zheng
Citations: 34
h-index: 2
Zhihan Li
Zhihan Li
Citations: 132
h-index: 6
Bohan Li
Bohan Li
Citations: 18
h-index: 1

통합된 음성 기반 모델은 언어 모델이 학습할 수 있을 뿐만 아니라 고품질 웨이브폼으로 디코딩될 수 있는 통합적인 토큰화 공간을 요구합니다. 그러나 기존의 음성 토크나이저는 종종 이러한 요구 사항을 동시에 만족시키지 못하여, 아키텍처 복잡성이 증가하고 훈련 설계가 더욱 복잡해지는 경향이 있습니다. 본 논문에서는 통합 생성-이해 모델링을 위해 설계된 연속적인 통합 음성 토큰화 모델인 HoliTok을 제안합니다. HoliTok은 48kHz 음성을 128차원 잠재 벡터 시퀀스로 압축하여, 25Hz 간격으로 표현합니다. 이 모델은 신호 수준의 충실도를 유지하고, 의미 정보를 통합하며, 강력한 잠재 학습 가능성을 갖도록 점진적인 전략을 통해 훈련됩니다. 이러한 토큰화를 기반으로, 음성 합성 및 인식에 사용되는 통합 AR+DiT 모델을 구축했습니다. 동일한 잠재 벡터 시퀀스가 생성 특화 작업과 통합된 생성-이해 작업 모두를 지원합니다. 실험 결과, HoliTok은 경쟁력 있는 재구성 정확도를 달성하고, 고품질의 제어 가능한 합성을 위한 생성 학습 능력을 향상시킵니다. 또한, 평가된 표현 방식 중 HoliTok만이 추가적인 최적화 트릭 없이 통합된 생성-이해 아키텍처에서 안정적으로 작동하는 것으로 나타났습니다. 이러한 결과는 HoliTok이 효과적인 음성 토크나이저이자 통합된 음성 언어 모델링을 위한 기초적인 표현 인터페이스로 활용될 수 있음을 시사합니다. 코드: https://github.com/bovod-sjtu/HoliTok.

Original Abstract

Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-quality waveforms. Existing speech tokenizers, however, often fail to satisfy these requirements simultaneously, leading to increased architectural complexity and more involved training designs. We propose HoliTok, a continuous Holistic speech Tokenization model designed for unified generation-understanding modeling. HoliTok encodes 48~kHz speech into a compact 25~Hz sequence of 128-dimensional latents. It is trained with a progressive strategy that jointly preserves signal-level fidelity, incorporates semantic information, and maintains strong latent learnability. Based on this tokenization, we build a unified AR+DiT model for speech synthesis and recognition, where the same latent sequence supports both generation-specific and unified generation-understanding tasks. Experiments show that HoliTok achieves competitive reconstruction fidelity, improves generative learnability for high-quality and controllable synthesis, and, among the evaluated representations, is the only one that operates robustly in our unified generation-understanding architecture without additional optimization tricks. These results suggest that HoliTok serves as an effective speech tokenizer and a foundational representation interface for unified spoken language modeling. The code is available at: https://github.com/bovod-sjtu/HoliTok.

3 Citations
0 Influential
40.540251005511 Altmetric
13.9 Score
Original PDF
14

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!