2606.25391v1 Jun 24, 2026 cs.SD

소리에서 장면으로: 대규모 오디오 언어 모델의 상황 인지 청각 장면 이해 평가를 위한 벤치마크

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models

Yutong Song
Yutong Song
Citations: 19
h-index: 2
Pengfei Zhang
Pengfei Zhang
Citations: 64
h-index: 3
Henry Peng Zou
Henry Peng Zou
University of Illinois Chicago
Citations: 827
h-index: 17
Honghui Xu
Honghui Xu
Citations: 5
h-index: 2
Amir M. Rahmani
Amir M. Rahmani
University of California Irvine
Citations: 617
h-index: 10
Hoang Nguyen
Hoang Nguyen
Citations: 185
h-index: 8
K. Sharif
K. Sharif
Citations: 68
h-index: 5
Wenjun Huang
Wenjun Huang
Citations: 7
h-index: 2
Pinxin Liu
Pinxin Liu
Citations: 43
h-index: 4

최근 개발된 대규모 오디오 언어 모델(LALM)은 음성, 소리, 음악을 포함한 다양한 음향 영역에서 뛰어난 성능을 보이고 있습니다. 그러나 기존의 벤치마크는 이러한 영역들을 개별적으로 평가하는 경향이 있으며, 현실 세계의 청각 장면에서 여러 음원들이 동시에 나타날 때 발생하는 복잡한 맥락적 관계를 간과합니다. 실제 환경에서의 청각 정보 이해에는 상황 인지 청각 장면 이해(CASU) 능력이 필요하며, 이는 다양한 음향 정보를 통합하여 전체적인 장면을 이해하는 능력입니다. 이러한 능력을 평가하기 위해, 우리는 CASU 벤치마크를 소개합니다. 이 벤치마크는 오디오 LLM이 음성, 음향 이벤트(예: 안내 방송), 배경 환경(예: 교통 소음) 등으로 구성된 청각 장면을 해석하고, 이러한 요소들 간의 논리적 관계에 대해 추론할 수 있는지 평가합니다. 우리는 실제 환경의 소리와 합성 음성을 결합하여 시간적으로 정확한, 준-인공적인 오디오 스트림을 구축하기 위한 확장 가능한 파이프라인을 제안했습니다. 이 데이터를 기반으로, 장면 이해 능력을 측정하기 위해 네 가지 과제를 설계했습니다. 이러한 과제는 맥락적 질문 응답, 장면 내 개체 추출, 화자 역할 추론 및 장면 조작을 통한 반사실적 추론을 포함합니다. 여러 LALM에 대한 실험 결과는 효과적인 청각 장면 이해를 위해서는 음성이나 소리 단독뿐만 아니라 모든 음향 영역의 통합이 필요하며, 이는 LALM에서 복잡한 오디오 이해 능력을 향상시키기 위한 CASU의 중요성을 강조한다는 것을 보여줍니다.

Original Abstract

Recent Large Audio Language Models (LALMs) have achieved remarkable progress in audio perceptual tasks across individual acoustic layers, including speech, sound, and music. However, existing benchmarks predominantly evaluate these layers in isolation, overlooking the complex contextual relationships that arise when multiple acoustic sources co-occur in real-world auditory scenes. Real-world auditory interpretation requires Context-Aware Auditory Scene Understanding (CASU): the ability to comprehend the holistic scene by integrating sound layers. To evaluate this capability, we introduce the CASU benchmark, which assesses whether Audio LLMs can interpret auditory scenes composed of speech, acoustic events (e.g., announcements), and background environments (e.g., traffic), and reason about the logical relationships between these layers. We propose a scalable pipeline for constructing time-accurate, semi-synthetic audio streams by composing real-world scene sounds with synthetic speech. Building on this data, we design four tasks that probe scene understanding: contextual question answering, entity extraction from the scene, speaker role inference, and counterfactual reasoning where scene is manipulated. Experiments across multiple LALMs demonstrate that effective auditory scene understanding requires integration over all auditory layers, rather than reliance on speech or sound alone, underscoring the necessity of CASU for advancing complex audio understanding in LALMs.

1 Citations
0 Influential
8.5 Altmetric
43.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!