2608.03264v1 Aug 04, 2026 cs.MM

듣고 보고: 오디오-비주얼 인스턴스 분할을 위한 상태 기반 청취 방식

Hear to See: Discerning Stateful Listening for Audio-Visual Instance Segmentation

Miao Zhang
Miao Zhang
Citations: 57
h-index: 4
Leiye Liu
Leiye Liu
Citations: 32
h-index: 2
Jiahong Jiang
Jiahong Jiang
Citations: 23
h-index: 2
Jingjing Li
Jingjing Li
Citations: 787
h-index: 11
Jialong Zhong
Jialong Zhong
Citations: 12
h-index: 2
Kai Peng
Kai Peng
Citations: 0
h-index: 0
Tingwei Liu
Tingwei Liu
Citations: 47
h-index: 3
Wei Ji
Wei Ji
Citations: 55
h-index: 3
Yongri Piao
Yongri Piao
Dalian University of Technology
Citations: 1,171
h-index: 16
Huchuan Lu
Huchuan Lu
Citations: 82
h-index: 5

오디오-비주얼 인스턴스 분할(AVIS)은 개별 소리 나는 객체를 정확하게 식별하고 픽셀 수준 마스크를 사용하여 추적하는 것을 요구합니다. 기존 방법들은 중첩된 음향 이벤트와 시각적 인스턴스를 연결하고 비동기적인 오디오-비주얼 동역학을 처리하는 데 어려움을 겪습니다. 따라서 다음과 같은 두 가지 중요한 질문이 제기됩니다: 모델은 어떻게 중첩된 소리 발생원과 시각적 인스턴스 간의 정확한 대응 관계를 설정할 수 있는가? 그리고 모델은 오디오 및 비주얼 신호가 시간적으로 불일치하는 경우에도 안정적인 추적을 유지할 수 있는가? 본 논문에서는 이러한 과제를 해결하기 위해 Hear to See (H2S)라는 새로운 방법을 제안합니다. Acoustic-Semantic Projector (ASP)는 혼합된 오디오를 분리하고 의미 영역에서 공간 영역으로의 계층적 대응 관계를 설정합니다. Asynchronous Dynamics Modulator (ADM)는 오디오에 의해 조절되는 Mamba를 사용하여 상태 전환을 적응적으로 조정하며, 동적인 변화 시 현재 정보를 우선시하고 안정적인 기간 동안에는 연속성을 유지합니다. AVISeg 데이터셋에서의 실험 결과, H2S는 SOTA(State-of-the-Art) 성능을 달성했으며, COCO 사전 훈련된 ResNet50 모델을 사용하여 48.54 mAP를 기록하며 이전 최고 성능보다 7.8% 향상되었습니다. 논문이 채택되면 코드가 공개될 예정이며, 소스 코드는 https://github.com/leiyeliu/H2S 에서 확인할 수 있습니다.

Original Abstract

Audio-visual instance segmentation (AVIS) requires accurately identifying and tracking individual sounding objects with pixel-level masks. Existing methods struggle to match overlapping acoustic events with visual instances and handle asynchronous audio-visual dynamics. Therefore, two critical questions arise: how can a model establish precise correspondence between overlapping sound sources and visual instances, and how can a model maintain robust tracking when audio and visual signals are temporally misaligned?This paper proposes Hear to See (H2S), addressing these challenges through two mechanisms. The Acoustic-Semantic Projector (ASP) disentangles mixed audio and establishes hierarchical correspondence from semantic to spatial domains. The Asynchronous Dynamics Modulator (ADM) adaptively adjusts state transitions via audio-modulated Mamba, prioritizing current information during dynamic variations and maintaining continuity in stable periods.Experiments on AVISeg show H2S achieves SOTA performance, attaining 48.54 mAP with a COCO pretrained ResNet50 and surpassing the previous by 7.8\%. The code will be open-sourced once the paper is accepted. The source code will be publicly available at https://github.com/leiyeliu/H2S.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!