2607.20166v1 Jul 22, 2026 cs.SD

Audio-Zero: 레이블 없이 스스로 발전하는 정밀 오디오 추론

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning

Baolong Bi
Baolong Bi
Citations: 600
h-index: 14
Siqian Tong
Siqian Tong
Citations: 15
h-index: 1
Xuan Li
Xuan Li
Citations: 16
h-index: 3
Yiwei Wang
Yiwei Wang
Citations: 490
h-index: 13
Yujun Cai
Yujun Cai
Citations: 21
h-index: 2
Shenghua Liu
Shenghua Liu
Citations: 489
h-index: 13
C. Hao
C. Hao
Citations: 50
h-index: 4
Chaozhuo Li
Chaozhuo Li
Citations: 364
h-index: 11

대규모 오디오 언어 모델(LALM)은 음향 이해 측면에서 빠르게 발전했지만, 여전히 세분화된 오디오 추론(예: 이벤트 순서, 반복 및 지속 시간 인식)에는 어려움을 겪고 있습니다. 기존의 사후 훈련 방법은 비용이 많이 드는 외부 레이블에 크게 의존하거나, 제한적인 수준의 의미 정보를 제공합니다. 이러한 격차를 해소하기 위해, 본 논문에서는 LALM 분야 최초의 레이블 없는 자기 진화 프레임워크인 Audio-Zero를 소개합니다. Audio-Zero는 정밀한 청각 인식 및 추론 능력을 향상시키기 위해, 레이블이 없는 오디오 대비 쌍을 이용하여 청각적 자가 학습 게임을 구성합니다. 대부분의 플레이어는 참조 오디오를 듣는 반면, 일부 플레이어는 미묘하게 변형된 오디오를 듣습니다. 모델은 먼저 자신이 들리는 내용을 설명하는 단서를 생성하고, 이어서 단서 간의 불일치를 추론하여 이상한 플레이어를 식별합니다. 이상한 플레이어는 설계에 의해 미리 알려져 있으므로, 이 게임은 어떠한 주석 처리된 답변 없이도 검증 가능한 보상을 제공합니다. Qwen2-Audio-7B-Instruct 및 Qwen2.5-Omni-7B를 사용하여 TREA, MMAU Test-mini 및 MMAR 데이터셋에서 수행한 실험 결과, Audio-Zero는 전반적인 오디오 이해 능력을 유지하면서 정밀한 오디오 추론 능력을 향상시키는 것을 확인했습니다. 또한, 진화적 분석과 진단 분석을 통해 게임의 압력으로 인해 점점 더 세분화된 청각 묘사가 자연스럽게 나타나는 것을 확인할 수 있었습니다.

Original Abstract

Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.g., recognizing event order, repetitions and duration). Existing post-training methods heavily rely on expensive external labels or provide only coarse semantic signals. To bridge this gap, we introduce Audio-Zero, the first label-free self-evolution framework in the field of LALMs that improves fine-grained auditory perception and reasoning. Audio-Zero constructs an auditory self-play game from unlabeled audio contrast pairs: most players hear a reference audio, while one odd listener hears a subtle variant. The model first generates clues describing what it hears and then identifies the odd listener by reasoning over inconsistencies among clues. Since the odd listener is known by construction, the game provides verifiable rewards without any annotated answers. Experiments with Qwen2-Audio-7B-Instruct and Qwen2.5-Omni-7B on TREA, MMAU Test-mini and MMAR show that Audio-Zero improves fine-grained audio reasoning while preserving broad audio understanding. Evolutionary and diagnostic analyses further reveal that increasingly fine-grained auditory descriptions emerge naturally from game pressure.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!