오디오 추론을 위한 진화하는 평가 기준 기반 강화 학습
Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning
오디오 추론은 기계가 음향 환경을 이해하는 데 필수적입니다. 검증 가능한 보상을 활용한 강화 학습은 이러한 추론 능력을 향상시킬 수 있지만, 기존의 보상 설계 방식은 한계점을 가지고 있습니다. 결과 기반 보상은 최종 답변만을 감독하며 모델이 오디오 내용을 고려하지 않고도 답에 도달하도록 만들 수 있고, 과정 기반 보상은 추론 자체를 평가하지만, 질문마다 적응하지 못하고 음향 증거에 기반하지 않은 거칠고 미리 정의된 기준에 의존합니다. 또한, 질문은 감각 인지 능력을 요구하는 것과 다단계 추론을 요구하는 것으로 다르며, 어떤 정적인 기준이라도 정책이 개선됨에 따라 그 효과가 약화됩니다. 따라서 미세하게 조정되고 음향 기반이며 적응적인 보상을 사용하여 추론 과정을 감독하는 것이 매우 중요하지만, 이러한 보상을 수동으로 설계하는 것은 모든 샘플에 대해 실용적이지 않습니다. 이에 우리는 AudioRubrics라는 강화 학습 프레임워크를 소개합니다. AudioRubrics는 자체적으로 진화하는 음향 기반 평가 기준을 사용하여 오디오 추론을 감독합니다. AudioRubrics는 원시 파형에서 샘플별 평가 기준을 합성하고, 모델의 자체 실행 결과를 조건으로 하여 각 그룹별로 기준을 재생성하고 재가중하여 지속적인 학습 신호를 제공함으로써 정적 기준이 포화될 때에도 현재 정책의 약점을 계속 타겟팅합니다. 세 가지 오디오 추론 벤치마크에 대한 종합적인 평가 결과, AudioRubrics는 다양한 오픈 소스 및 학습 기반 모델보다 훨씬 뛰어난 성능을 보였습니다. 또한, 분석 결과 평가 기준 생성기와 판단기의 능력이 향상될수록 성능이 향상되며, AudioRubrics는 퇴화 또는 무한 성장을 피하는 안정적인 추론 길이에 수렴합니다. 음향 인지 능력의 개선은 음향 증거에 기반한 감독 방식의 효과를 더욱 입증합니다. 프로젝트 페이지는 https://audiorubrics.github.io 에서 확인할 수 있습니다.
Audio reasoning is essential for machine understanding of the acoustic world. Reinforcement learning with verifiable rewards can elicit such reasoning, yet existing reward designs are complementary in their limitations: outcome-based rewards supervise only the final answer and let the model reach it without attending to the audio, whereas process-based rewards score the reasoning itself but rely on coarse, hand-crafted, and fixed criteria that neither adapt to each question nor stay grounded in the acoustic evidence. Moreover, questions differ in what they demand, with some hinging on perception and others on multi-step reasoning, and any static criterion weakens as the policy improves. Supervising the reasoning process with fine-grained, audio-grounded, and adaptive rewards is therefore crucial, yet challenging since such rewards are impractical to design by hand for every sample. To this end, we introduce AudioRubrics, a reinforcement learning framework that supervises audio reasoning with self-evolving, audio-grounded rubric rewards. AudioRubrics synthesizes per-sample rubrics from the raw waveform and, conditioned on the model's own rollouts, regenerates and reweights criteria per group, supplying a continuous learning signal that keeps targeting the current policy's weaknesses as static criteria saturate. Comprehensive evaluations across three audio reasoning benchmarks reveal that AudioRubrics substantially outperforms a wide range of open-source and training-based baselines. Furthermore, our analysis shows that the gains scale with the capability of the rubric generator and judge, and AudioRubrics converges to a stable reasoning length that avoids both degenerate collapse and unbounded growth. The improvement in audio perception further demonstrates the effectiveness of anchoring supervision in the acoustic evidence. Our project page is available at https://audiorubrics.github.io.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.