MusTBENCH: 음악 LLM의 시간적 정렬 평가 및 발전
MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs
최근 개발된 대규모 오디오-언어 모델(LALMs)은 음악 콘텐츠 이해 능력에서 상당한 잠재력을 보여주었습니다. 그러나 이러한 모델들이 오디오의 정확한 시간 영역에 기반하여 응답하는지 여부는 아직 충분히 연구되지 않았습니다. 특히 음악 이해에서는 악기 등장 및 리듬 변화와 같은 시간적으로 국지화된 이벤트 형태로 중요한 정보가 자주 나타나므로, 이러한 제한점은 매우 중요합니다. 이 문제점을 해결하기 위해, 우리는 LALM의 시간적 정렬을 평가하기 위한 다섯 가지 시간 기반 질의응답 태스크로 구성된 음악 전문가 검증 벤치마크인 MusTBENCH를 소개합니다. 또한, 기존 모델의 시간적 정렬 능력을 향상시키기 위해, 음악 인코더 적응, LLM 적응, LLM 지도 학습 및 강화학습 기반 최적화를 포함하는 네 단계로 구성된 새로운 시간 최적화 방법인 MusT를 제안합니다. MusTBENCH에서의 실험 결과는 기존 LALM이 정확한 시간적 정렬에 어려움을 겪고 있으며, MusT가 강력한 기준 모델보다 상당한 성능 향상을 가져온다는 것을 보여줍니다. 이러한 결과는 현재의 LALM에서 시간적 정렬이 중요한 누락된 기능임을 입증하며, MusTBENCH를 시간 기반 음악 이해 분야의 미래 연구를 위한 도전적인 벤치마크로 자리매김합니다.
Recent Large Audio-Language Models (LALMs) have demonstrated promising abilities in understanding musical content. However, whether their responses are grounded in the correct temporal regions of the audio remains underexplored. This limitation is particularly critical for music understanding, where key information often occurs as temporally localized events, such as instrument entries and rhythmic transitions. To address this gap, we introduce MusTBENCH, a music-expert-validated benchmark designed to evaluate temporal grounding in LALMs through five temporally grounded question-answering tasks. To further improve temporal grounding in existing models, we propose MusT, a novel four-stage temporal optimization recipe spanning music encoder adaptation, LLM adaptation, LLM supervised fine-tuning, and RL-based optimization. Experiments on MusTBENCH show that existing LALMs struggle with precise temporal grounding, while MusT brings significant improvements over strong baselines. These results establish temporal grounding as a key missing capability in current LALMs and position MusTBENCH as a challenging benchmark for future research in temporally grounded music understanding.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.