2607.28896v1 Jul 30, 2026 cs.SD

TORUS: 통합 오디오 모델의 렌더링 이해 및 자기 일관성 검증

TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models

Abhishek Mukherji
Abhishek Mukherji
Citations: 8
h-index: 2
Dinesh Manocha
Dinesh Manocha
Citations: 41
h-index: 4
Aryan Vijay Bhosale
Aryan Vijay Bhosale
Citations: 0
h-index: 0
Harshit Rajgarhia
Harshit Rajgarhia
Citations: 3
h-index: 1

오디오 이해, 오디오 생성, 그리고 점점 더 많이 사용되는 오디오 편집 기능을 제공하는 통합 오디오 모델이 빠르게 증가하고 있습니다. 하지만 이러한 모델에 대한 기본적인 질문은 여전히 답을 찾지 못했습니다: 통합 모델의 두 가지 구성 요소는 동일한 오디오 데이터에 대해 일치하는 의견을 가지고 있을까요? 현재까지, 각 기능은 특수화된 벤치마크에서 개별적으로 평가되며, 모델이 자신의 생성 결과에 대해 이해할 수 있는지 여부는 평가되지 않습니다. 본 논문에서는 오디오 전용 통합 모델의 자기 일관성을 검증하기 위한 첫 번째 테스트 도구인 TORUS를 제시합니다. TORUS는 음성, 소리 및 음악을 아우르는 5가지 작업 유형에서 수행되는 432개의 6지선다형 질문으로 구성된 48개의 3단계 자기 일관성 테스트로 이루어져 있습니다. 우리는 다섯 가지 공개 통합 모델과 최첨단 전문 생성, 편집 및 이해 모델을 결합한 Cascade Baseline을 종합적으로 평가했습니다. 가장 우수한 성능을 보인 통합 모델은 전체 질문의 50.5%를 맞춘 반면, Cascade Baseline은 63.2%, 그리고 단순히 무작위로 추론했을 때의 확률은 16.7%였습니다. 모델들은 특히 오디오 편집 작업에서 어려움을 겪는 것으로 나타났습니다. 평가된 오디오 모델(전문 및 통합) 전반에 걸쳐 제한적인 자기 일관성이 관찰되었으며, 따라서 자기 일관성은 향후 오디오 시스템을 평가하는 데 필수적인 요소로 간주됩니다.

Original Abstract

Unified audio models capable of audio understanding, audio generation and, increasingly, audio editing are proliferating rapidly. Yet a basic question about them remains unanswered: do the two heads of a unified model agree about the same audio? Current practice evaluates each capability in isolation on specialized benchmarks, and never asks whether a model can make sense of its own generations. We present TORUS, the first self-coherence test for audio-native unified models. TORUS comprises 48 three-stage self-coherence tests carrying 432 six-option questions spanning speech, sound and music across five task families. We holistically evaluate five open unified models alongside a Cascaded Baseline that combines state-of-the-art specialized generation, editing and understanding models. The best unified model answers 50.5% of questions against the Cascaded Baseline's 63.2% and a 16.7% chance floor. Models struggle on audio editing. Among the evaluated audio models (specialized and unified), we observe limited self-coherence, and thus position self-coherence as an essential test for future audio systems.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!