TORUS: 통합 오디오 모델의 렌더링 이해 및 자기 일관성 검증
TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models
오디오 이해, 오디오 생성, 그리고 점점 더 많이 사용되는 오디오 편집 기능을 제공하는 통합 오디오 모델이 빠르게 증가하고 있습니다. 하지만 이러한 모델에 대한 기본적인 질문은 여전히 답을 찾지 못했습니다: 통합 모델의 두 가지 구성 요소는 동일한 오디오 데이터에 대해 일치하는 의견을 가지고 있을까요? 현재까지, 각 기능은 특수화된 벤치마크에서 개별적으로 평가되며, 모델이 자신의 생성 결과에 대해 이해할 수 있는지 여부는 평가되지 않습니다. 본 논문에서는 오디오 전용 통합 모델의 자기 일관성을 검증하기 위한 첫 번째 테스트 도구인 TORUS를 제시합니다. TORUS는 음성, 소리 및 음악을 아우르는 5가지 작업 유형에서 수행되는 432개의 6지선다형 질문으로 구성된 48개의 3단계 자기 일관성 테스트로 이루어져 있습니다. 우리는 다섯 가지 공개 통합 모델과 최첨단 전문 생성, 편집 및 이해 모델을 결합한 Cascade Baseline을 종합적으로 평가했습니다. 가장 우수한 성능을 보인 통합 모델은 전체 질문의 50.5%를 맞춘 반면, Cascade Baseline은 63.2%, 그리고 단순히 무작위로 추론했을 때의 확률은 16.7%였습니다. 모델들은 특히 오디오 편집 작업에서 어려움을 겪는 것으로 나타났습니다. 평가된 오디오 모델(전문 및 통합) 전반에 걸쳐 제한적인 자기 일관성이 관찰되었으며, 따라서 자기 일관성은 향후 오디오 시스템을 평가하는 데 필수적인 요소로 간주됩니다.
Unified audio models capable of audio understanding, audio generation and, increasingly, audio editing are proliferating rapidly. Yet a basic question about them remains unanswered: do the two heads of a unified model agree about the same audio? Current practice evaluates each capability in isolation on specialized benchmarks, and never asks whether a model can make sense of its own generations. We present TORUS, the first self-coherence test for audio-native unified models. TORUS comprises 48 three-stage self-coherence tests carrying 432 six-option questions spanning speech, sound and music across five task families. We holistically evaluate five open unified models alongside a Cascaded Baseline that combines state-of-the-art specialized generation, editing and understanding models. The best unified model answers 50.5% of questions against the Cascaded Baseline's 63.2% and a 16.7% chance floor. Models struggle on audio editing. Among the evaluated audio models (specialized and unified), we observe limited self-coherence, and thus position self-coherence as an essential test for future audio systems.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.