2605.27840v1 May 27, 2026 eess.AS

LoSATok: 교차 도메인 오디오 이해 및 생성을 위한 저차원 의미-음향 토크나이저

LoSATok: Low-dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation

Guoyang Zeng
Guoyang Zeng
Citations: 40
h-index: 2
Zhisheng Zhang
Zhisheng Zhang
Citations: 33
h-index: 3
Xiang Li
Xiang Li
Citations: 37
h-index: 3
Yixuan Zhou
Yixuan Zhou
Citations: 206
h-index: 9
Jing Peng
Jing Peng
Citations: 9
h-index: 1
Zhiyong Wu
Zhiyong Wu
Citations: 194
h-index: 4

오디오 토크나이저는 오디오 이해와 생성을 통합하는 데 필수적인 요소입니다. 이해는 고수준의 의미를 필요로 하는 반면, 생성은 의미와 음향 세부 사항을 요구합니다. 기존의 통합 토크나이저는 이러한 모든 정보를 고차원 연속 공간에 함께 인코딩하여, 생성 모델인 디퓨전 트랜스포머(DiT)의 모델링 부담을 증가시킵니다. 본 연구에서는 교차 도메인 오디오 이해 및 생성을 위한 저차원 오디오 토크나이저인 LoSATok을 제안합니다. 1280차원의 의미 인코더 특징은 압축 가능하다는 관찰에 따라, 우리는 시간적 특징 일관성을 유지하기 위해 제안된 시간 관계 손실로 규제되는 Semantic Bottleneck(SemBo)을 도입하여 이를 128차원으로 압축합니다. 또한, 고차원 및 저차원 의미 신호를 모두 활용하는 이중 레벨의 의미 지도 방법을 설계하여, 토크나이저가 간결한 잠재 공간 내에서 의미와 음향 세부 사항을 동시에 포착할 수 있도록 합니다. 음성, 음악 및 일반 오디오 데이터에 대한 실험 결과, SemBo는 강력한 저차원 의미 표현 능력을 유지하며, LoSATok은 여러 의미 표현 방식과 비교하여 경쟁력 있는 이해 성능을 보입니다. 또한, LoSATok은 음성, 음악 및 오디오 생성에서의 DiT 모델링 성능을 지속적으로 향상시킵니다. 이러한 결과는 LoSATok의 저차원 표현이 오디오 이해와 생성을 효과적으로 지원할 수 있음을 보여줍니다. 저희 코드는 다음 GitHub 저장소에서 확인할 수 있습니다: https://github.com/wxzyd123/LoSATok.

Original Abstract

Audio tokenizers are fundamental to unifying audio understanding and generation. Understanding requires high-level semantics, while generation demands semantic and acoustic details. Existing unified tokenizers jointly encode both in high-dimensional continuous latents, which increases the modeling burden of Diffusion Transformers (DiTs) for generation. We propose LoSATok, a low-dimensional audio tokenizer for cross-domain audio understanding and generation. Motivated by the observation that 1280-dimensional semantic encoder features are compressible, we introduce a Semantic Bottleneck that compresses them into 128 dimensions, regularized by the proposed time-relation loss for temporal feature consistency. We further design a dual-level semantic supervision method that leverages both high- and low-dimensional semantic signals, enabling the tokenizer to jointly capture semantics and acoustic details within a compact latent space. Experiments on speech, music, and general audio show that SemBo preserves strong low-dimensional semantic capacity and LoSATok retains competitive understanding performance compared with several semantic representations, while consistently improving DiT modeling performance on speech, music, and audio generation. These results demonstrate that LoSATok's low-dimensional representations can effectively support audio understanding and generation. Our code is provided at https://github.com/wxzyd123/LoSATok.

1 Citations
0 Influential
36.92453324894 Altmetric
6.9 Score
Original PDF
11

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!