2608.04479v1 Aug 05, 2026 cs.SD

AudioScape-TTA: 세밀한 텍스트 음성 변환 평가를 위한 구조화된 음향 환경 벤치마크

AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation

Xiaoda Yang
Xiaoda Yang
Citations: 514
h-index: 10
Yuguang Yang
Yuguang Yang
Citations: 888
h-index: 11
Yan Rong
Yan Rong
Citations: 119
h-index: 7
Shan Yang
Shan Yang
Citations: 37
h-index: 3
Jinting Wang
Jinting Wang
Citations: 55
h-index: 4
Shengyu Li
Shengyu Li
Citations: 0
h-index: 0
Li Liu
Li Liu
Citations: 5
h-index: 1

최근 텍스트 음성 변환(TTA) 기술은 자연어 설명을 기반으로 사실적인 오디오를 생성하는 데 상당한 발전을 이루었습니다. 그러나 생성된 오디오가 복잡한 텍스트 지시사항을 얼마나 충실하게 반영하는지 판단하는 것은 여전히 어려운 과제입니다. 기존의 벤치마크는 주로 전체적인 유사성 측정에 의존하며, 세밀한 의미 오류에 대한 제한적인 정보를 제공합니다. 이러한 한계를 극복하기 위해, 본 연구에서는 세밀한 TTA 평가를 위한 구조화되고 복잡성을 고려한 벤치마크인 extbf{AudioScape-TTA}를 제안합니다. AudioScape-TTA는 모달리티 인식 의미 구조를 통해 실제적인 음향 환경을 표현하고, 이벤트 밀도와 구조적 복잡성을 사용하여 생성 복잡도를 특징짓습니다. 이러한 어노테이션을 기반으로, 이벤트 구현, 음향 속성 및 음성 내용의 세밀한 의미 기준을 통한 오디오 기반 평가 프레임워크를 제안합니다. 본 벤치마크는 2,258개의 오디오-텍스트 쌍과 25,707개의 이진 질문 답변 항목(QA rubric)으로 구성되어 있으며, 이를 통해 TTA 시스템에 대한 확장 가능하고 해석 가능한 분석을 제공합니다. 13개의 대표적인 오픈 소스 TTA 모델에 대한 실험 결과, 세밀한 속성 제어, 음성 내용 보존 및 복합적인 음향 환경 생성 측면에서 지속적인 한계가 존재하는 것으로 나타났습니다. 인간 평가를 통해 본 연구에서 제안하는 질문 답변 기반 평가는 기존의 전체 유사성 측정 방법보다 인간의 의미 판단과 더 높은 일관성을 보이는 것을 확인했습니다.

Original Abstract

Text-to-audio (TTA) generation has recently achieved remarkable progress in synthesizing realistic audio from natural language descriptions. However, determining whether generated audio faithfully satisfies complex textual instructions remains challenging. Existing benchmarks mainly rely on global similarity metrics, providing limited insight into fine-grained semantic failures. To address this limitation, we introduce \textbf{AudioScape-TTA}, a structured and complexity-aware benchmark for fine-grained TTA evaluation. AudioScape-TTA represents realistic soundscapes through modality-aware semantic structures and characterizes generation complexity using event density and structural complexity. Based on these annotations, we propose a rubric-based audio-grounded evaluation framework that verifies event realization, acoustic attributes, and speech content through fine-grained semantic criteria. The benchmark contains 2,258 audio-text pairs with 25,707 binary QA rubrics, enabling scalable and interpretable analysis of TTA systems. Experiments on 13 representative open-source TTA models reveal persistent limitations in fine-grained attribute control, speech-content preservation, and compositional soundscape generation. Human validation further demonstrates that our rubric-based evaluation achieves stronger alignment with human semantic judgments than conventional global similarity metrics.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!