SelectTSL: 프롬프트 기반 선택적 대상 음원 위치 추정 기술 - 복잡한 환경에서의 응용
SelectTSL: Prompt-Guided Selective Target Sound Localization in Complex Scenarios
사람은 복잡한 환경에서 특정 대상 음원에 집중하여 그 방향을 추정할 수 있지만, 현재의 딥러닝 기반 시스템에서는 이러한 선택적인 위치 추정이 여전히 어려운 과제입니다. 심층 학습을 활용한 음원 위치 추정 기술은 상당한 발전을 이루었지만, 대부분의 방법은 모든 활성 음원을 구분 없이 위치 추적합니다. 반면, 대상 음원 추출(TSE) 기술은 다중 모드 프롬프트를 사용하여 음원을 추출하지만, 정확한 위치 추정에 필요한 다채널 공간 정보를 제대로 보존하지 못하는 경우가 많습니다. 이러한 격차를 해소하기 위해, 본 연구에서는 프롬프트 기반의 선택적 대상 음원 위치 추정 문제를 정의하고, SelectTSL이라는 통합 아키텍처를 제안합니다. SelectTSL은 다중 음원이 존재하는 환경에서 사용자가 지정한 특정 대상 음원의 위치만 추정합니다. 구체적으로, 우리는 대상에 대한 정보를 활용하는 선택적인 위치 추정 전략을 설계했습니다. 이 전략은 프롬프트 기반의 선택적 어텐션 모듈(PGSA)을 사용하여 프롬프트 정보가 반영된 임베딩을 생성하며, 이러한 임베딩은 채널 간 위상차(IPD) 향상기로 전달되어 원본 위상 정보를 정제하고, 대상 음원의 크기와 결합하여 도래 방향(DoA)과 대상 음원 개수를 동시에 추정합니다. 이러한 통합 설계는 사용자가 지정한 대상 음원의 공간 정보를 효과적으로 활용하여 선택적인 위치 추정을 가능하게 하며, 또한 시간에 따라 변하는 대상 음원 개수에도 대응할 수 있습니다. 합성 데이터와 실제 녹음 데이터를 사용하여 수행한 광범위한 실험 결과, 제안된 방법은 기존의 다른 방법들을 능가하는 성능을 보였으며, 실제 음향 환경에서도 뛰어난 일반화 성능을 나타냈습니다.
Humans can selectively attend to a target sound and estimate its direction in complex scenarios, whereas such selective localization remains challenging for current deep learning-based systems. Sound source localization (SSL) has achieved remarkable success with deep learning, yet most methods localize all active sources without selectivity. Conversely, target sound extraction (TSE) extracts sources using multimodal prompts but typically fails to preserve the multichannel spatial information required for accurate localization. To bridge this gap, we formulate the task of prompt-guided selective target sound localization and propose SelectTSL, an end-to-end architecture that localizes only the user-specified target in multi-source acoustic scenes. Specifically, we design a target-aware selective localization strategy that employs a Prompt-Guided Selective Attention Module (PGSA) to generate prompt-informed embeddings. These embeddings guide an inter-channel phase difference (IPD) enhancer to refine raw phase cues, fusing with target magnitudes to jointly estimate direction of arrival (DoA) and target-source cardinality, i.e., the number of target sound sources. This coupled design effectively focuses on the user-specified target spatial cues for selective localization and also handles time-varying numbers of target sources. Extensive experiments on both synthetic data and real-world recordings demonstrate that our proposed method consistently outperforms other baselines and exhibits robust generalization to real acoustic environments.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.