공간-옴니: FOA 인코딩을 통한 다중 모드 LLM에서 공간 오디오 이해 통합
Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding
최근의 다중 모드 대규모 언어 모델(LLM)은 주로 오디오를 단일 채널 신호로 처리하여, 음원 위치 파악, 공간 관계 추론 및 공간 장면 이해에 필요한 공간 오디오의 공간 정보를 활용하지 못합니다. 본 연구에서는 Spatial-Omni라는 경량화된 방법을 제안합니다. 이 방법은 SO-Encoder를 사용하여 1차 방향성 음향(FOA) 공간 오디오를 기존 Omni LLM에 독립적인 모달리티로 통합하며, 기존 오디오 인코더를 수정하지 않습니다. SO-Encoder는 제한적인 추가적인 계산 비용으로 공간 토큰을 제공하고, 효율적인 단계별 학습을 통해 공간 오디오 이해 능력을 향상시킵니다. 학습 및 평가를 지원하기 위해, 공개 데이터, 실제 녹음 및 시뮬레이션을 활용하여 40만 개의 FOA 공간 오디오 클립과 210만 개의 공간 질의응답 쌍으로 구성된 SO-Dataset, SO-QA 및 SO-Bench를 구축했습니다. SO-Bench는 기본적인 탐지 및 위치 추정, 공간 관계 이해, 복잡한 공간 추론을 포함하여 16가지의 공간 오디오 이해 하위 작업을 다룹니다. 실험 결과, Spatial-Omni는 기존의 공개된 대규모 오디오-언어 모델(LALM)과 Omni LLM 모델보다 공간 오디오 이해 작업에서 우수한 성능을 보였으며, 일반적인 오디오 이해 능력 또한 상당한 수준으로 유지했습니다. 코드 및 데이터는 https://github.com/dieKarotte/Spatial-Omni 에서 확인할 수 있습니다.
Recent multimodal large language models mainly process audio as monaural signals, thereby discarding the spatial cues contained in spatial audio for sound localization, spatial relation reasoning, and spatial scene understanding. We propose Spatial-Omni, a lightweight method that implements SO-Encoder to inject First-Order Ambisonics (FOA) spatial audio into existing Omni LLMs as an independent modality, without modifying their original audio encoders. SO-Encoder provides spatial tokens with limited additional context cost and improves spatial audio understanding through efficient staged training. To support training and evaluation, we construct SO-Dataset, SO-QA, and SO-Bench from open-source data, real recordings, and simulations, containing 400K FOA spatial audio clips and 2.1M spatial question answering pairs. SO-Bench covers 16 spatial audio understanding subtasks, including basic detection and location estimation, spatial relation understanding, and complex spatial reasoning. Experiments show that Spatial-Omni outperforms existing open-source Large Audio-Language Models (LALMs) and Omni LLM models on spatial audio understanding tasks while retaining a reasonable level of general audio understanding. Code and data are available at https://github.com/dieKarotte/Spatial-Omni.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.