2606.05931v1 Jun 04, 2026 cs.CL

다중 모드 활용 여부: 능동적 모드 감지를 통한 쿼리 적응형 오디오-비주얼 인물 검색

To Be Multimodal or Not to Be: Query-Adaptive Audio-Visual Person Retrieval via Active Modality Detection

Mengjie Qian
Mengjie Qian
Citations: 175
h-index: 7
Mark J. F. Gales
Mark J. F. Gales
Citations: 137
h-index: 6
Kate Knill
Kate Knill
Citations: 130
h-index: 5
Erfan Loweimi
Erfan Loweimi
University of Cambridge
Citations: 597
h-index: 15
C. Chan
C. Chan
Citations: 0
h-index: 0
Abbas Haider
Abbas Haider
Citations: 13
h-index: 2
M. Awan
M. Awan
Citations: 0
h-index: 0
Joseph Kittler
Joseph Kittler
Citations: 67
h-index: 4
Hui Wang
Hui Wang
Citations: 224
h-index: 10
Guanfeng Wu
Guanfeng Wu
Citations: 13
h-index: 2

음성과 얼굴 정보를 이용하여 비디오 아카이브에서 특정 인물을 검색할 때, 시스템이 다중 모드를 활용해야 할까요? 실제 방송 아카이브에서는 벤치마크와 달리, 대상이 들리지만 보이지 않거나, 보이지만 들리지 않거나, 또는 둘 다인 경우가 존재합니다. 존재하지 않는 모드의 점수를 통합하면 노이즈가 발생하여 정확도가 최적의 단일 모드 시스템보다 떨어집니다. 우리는 능동적 모드를 감지하는 쿼리 적응형 프레임워크를 제안합니다. 이 프레임워크는 교차 모드 점수 일관성을 통해 작동하며, 두 모드가 모두 활성 상태일 때, 한 모드로 검색된 파일은 다른 모드에서도 높은 점수를 받습니다. 이러한 일관성은 모드가 없을 때 깨어집니다. 이러한 교차 모드 특징을 기반으로 하는 분류기는 89%의 감지 정확도를 달성합니다. BBC Rewind 코퍼스(12,000개 이상의 방송 비디오 포함)에서, 제안하는 적응형 시스템은 94.2%의 P@1 값을 얻어, 음성만 사용한 경우(82.9%), 얼굴만 사용한 경우(93.4%), 그리고 고정된 통합 방식(90.0%)보다 우수한 성능을 보입니다. 또한, 정답 모드 레이블이 있는 이상적인 시스템(96.6%)과의 격차를 64%나 줄였습니다.

Original Abstract

When retrieving a person from a video archive by voice and face, should the system be multimodal or not? In real-world broadcast archives, unlike curated benchmarks, a target may be heard but unseen, seen but unheard, or both. Fusing scores from an absent modality injects noise, degrading precision below the best unimodal system. We propose a query-adaptive framework that detects active modalities via cross-modal score consistency: when both modalities are active, files retrieved by one also score highly on the other; this agreement breaks down when a modality is absent. Classifiers driven by these cross-modal features achieve 89% detection accuracy. On the BBC Rewind corpus (with over 12,000 broadcast videos) the adaptive system attains 94.2% P@1, outperforming speaker-only (82.9%), face-only (93.4%), and fixed fusion (90.0%), recovering 64% of the gap to an oracle with ground-truth modality labels (96.6%).

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!