2608.03999v1 Aug 04, 2026 cs.SD

Agogic: LLM 기반 텍스트-음악 생성 시스템을 위한 성능 시간 기반 음악 토큰

Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Jingjia Mao
Jingjia Mao
Citations: 0
h-index: 0
Ruocheng Wu
Ruocheng Wu
Citations: 0
h-index: 0
Liaoyuan Fan
Liaoyuan Fan
Citations: 25
h-index: 4
M. Gao
M. Gao
Citations: 88
h-index: 6

텍스트-음악 언어 모델은 일반적으로 음악 토큰화 방식을 기본적으로 선택합니다. 이 방식은 모델의 핵심 구조, 데이터, 학습 방법에 밀접하게 연관되어 있어, 그 영향이 독립적으로 측정된 적은 없습니다. 본 연구에서는 사전 학습된 Qwen3.5 (0.8B-27B) 모델, 데이터셋, 예산 및 디코딩 방식을 고정하고, 7가지 토큰화 방식을 사용하여 각 방식의 성능 지표 변화를 분석했습니다. 결과는 예상외로 흥미로운 순서를 보여주었습니다. 즉, 모델 크기가 아닌, 토큰화 방식이 분포적 충실도에 가장 큰 영향을 미치는 변수라는 것을 확인했습니다. 모델 핵심 구조를 34배 확장해도 Frechet Music Distance (FMD) 값이 거의 변화하지 않는 반면, 토큰화 방식을 변경하면 FMD 값이 절반으로 줄어들었습니다. 본 연구에서 개발한 성능 기반의 고해상도 스트림(PMT: 10ms 단위 시간 정보, 음표별 강세, 멀티 트랙 구조; 609개의 심볼)은 0.8B 모델로 FMD 값 159를 달성했으며, 이는 동일 조건에서 사용된 비트 그리드 방식(FMD 272-286)보다 훨씬 우수한 성능입니다. 이러한 결과는 새로운 토큰화 방식을 적용한 26M 규모의 모델에서도 재현되었으며, 또한 다른 성능 기반 토크나이저에서도 관찰되어, 특정 어휘집이 아닌 해당 클래스의 일반적인 특성임을 시사합니다. PMT의 시작 지점을 비트 그리드 해상도로 조정하더라도 FMD 값에서 상당한 차이가 유지되어(n=500), 이는 단순한 해상도 문제로 인한 결과가 아님을 보여줍니다. 이러한 성능 향상은 분포적 측면에서 나타나며, 실제 청취 가능성 여부는 별도의 연구를 통해 검증할 예정입니다 (인간 평가 연구는 사전 등록 완료). 또한, 텍스트 설명과의 일관성은 낮지만, 디코딩 과정에서 간단한 제약을 추가하면 악기 F1 점수 (.28 -> .60) 및 정확한 키 인식률 (.16 -> .35)이 향상될 수 있습니다. 본 연구에서는 개발 도구, 25개 이상의 체크포인트 모델, 두 개의 데이터셋 (캡션/MIDI/ABC/오디오 정렬, 86.6k; 캡션이 포함된 데이터, 6.25M, 음악 분야에서 가장 큰 규모), 그리고 성능 진단 도구를 공개합니다. 기존의 텍스트-MIDI 변환 시스템은 학습 데이터 분포를 거의 동일하게 유지하며, 이는 캡션 내용과 관계없이 나타납니다 (분리된 데이터셋에서 코드 시간 일치율: 72% vs. 71%). 이제 음악 분야의 새로운 토큰화 방식 주장은 단순히 주장하는 것이 아니라, 객관적으로 측정될 수 있게 되었습니다.

Original Abstract

Text-to-music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe, its effect has never been measured in isolation. We fix pretrained Qwen3.5 (0.8B-27B), data, budget, and decoding, and swap only the representation across seven tokenizations, anchoring texture metrics to each representation's model-free ceiling. The ordering is clean and surprising: representation, not model size, is the binding variable for distributional fidelity. Scaling the backbone 34x barely moves Frechet Music Distance (FMD), whereas switching representation halves it. PMT, a performance-resolution stream we release (10 ms timing, per-note velocity, multi-track texture; 609 symbols), reaches FMD 159 at 0.8B against 272-286 for beat grids (1.7-1.8x lower, up to 2.8x elsewhere; non-overlapping bootstrap CIs), so a 0.8B performance-resolution model beats a 27B beat grid. It reappears on a 26M from-scratch backbone and a second performance-resolution tokenizer: a property of the class, not one lucky vocabulary. Nor is it a finer-lattice artifact: snapping PMT's onsets to the beat grids' resolution still leaves it 67-129 FMD ahead of both (n=500). The effect is distributional; whether it is audible is a separate question, left open by our probe, with a human study pre-registered. Native caption adherence is weak but separable: a lightweight decode-time constraint doubles instrument-F1 (.28 to .60) and Correct-Key (.16 to .35) at no distributional cost. We release the harness, 25+ checkpoints, two corpora (86.6k aligned across caption/MIDI/ABC/audio; 6.25M captioned, the largest for music), and an imprinting diagnostic: published text-to-MIDI systems reproduce their training distribution near-invariant to the caption (72% vs. 71% chord-time on disjoint domains). The field's next representation claim can now be measured, not asserted.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!