다중 모드 활용 여부: 능동적 모드 감지를 통한 쿼리 적응형 오디오-비주얼 인물 검색
To Be Multimodal or Not to Be: Query-Adaptive Audio-Visual Person Retrieval via Active Modality Detection
음성과 얼굴 정보를 이용하여 비디오 아카이브에서 특정 인물을 검색할 때, 시스템이 다중 모드를 활용해야 할까요? 실제 방송 아카이브에서는 벤치마크와 달리, 대상이 들리지만 보이지 않거나, 보이지만 들리지 않거나, 또는 둘 다인 경우가 존재합니다. 존재하지 않는 모드의 점수를 통합하면 노이즈가 발생하여 정확도가 최적의 단일 모드 시스템보다 떨어집니다. 우리는 능동적 모드를 감지하는 쿼리 적응형 프레임워크를 제안합니다. 이 프레임워크는 교차 모드 점수 일관성을 통해 작동하며, 두 모드가 모두 활성 상태일 때, 한 모드로 검색된 파일은 다른 모드에서도 높은 점수를 받습니다. 이러한 일관성은 모드가 없을 때 깨어집니다. 이러한 교차 모드 특징을 기반으로 하는 분류기는 89%의 감지 정확도를 달성합니다. BBC Rewind 코퍼스(12,000개 이상의 방송 비디오 포함)에서, 제안하는 적응형 시스템은 94.2%의 P@1 값을 얻어, 음성만 사용한 경우(82.9%), 얼굴만 사용한 경우(93.4%), 그리고 고정된 통합 방식(90.0%)보다 우수한 성능을 보입니다. 또한, 정답 모드 레이블이 있는 이상적인 시스템(96.6%)과의 격차를 64%나 줄였습니다.
When retrieving a person from a video archive by voice and face, should the system be multimodal or not? In real-world broadcast archives, unlike curated benchmarks, a target may be heard but unseen, seen but unheard, or both. Fusing scores from an absent modality injects noise, degrading precision below the best unimodal system. We propose a query-adaptive framework that detects active modalities via cross-modal score consistency: when both modalities are active, files retrieved by one also score highly on the other; this agreement breaks down when a modality is absent. Classifiers driven by these cross-modal features achieve 89% detection accuracy. On the BBC Rewind corpus (with over 12,000 broadcast videos) the adaptive system attains 94.2% P@1, outperforming speaker-only (82.9%), face-only (93.4%), and fixed fusion (90.0%), recovering 64% of the gap to an oracle with ground-truth modality labels (96.6%).
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.