SceneBind: 비전, 오디오 및 언어를 통해 무엇과 어디를 연결하는 방법
SceneBind: Binding What and Where Across Vision, Audio and Language
본 논문에서는 SceneBind를 제안합니다. SceneBind는 현실적인 장면을 전반적인 의미론적 정보와 3차원 공간 정보를 결합하여 표현하며, 이는 시각, 음성 및 언어 데이터를 포괄합니다. 기존의 다중 모달 인코더는 개별 객체의 의미론적 정보(즉, 무엇이 존재하는가)를 잘 파악하지만, 종종 명시적인 공간 구조(즉, 어디에 있는가)에 대한 이해가 부족합니다. SceneBind는 이러한 간극을 메우기 위해 각 장면을 전역 의미론적 임베딩과 객체 중심의 의미론적-공간적 슬롯을 결합한 의미론적-공간적 개체로 표현합니다. 이 표현은 객체 수준의 의미, 공간 속성 및 불확실성을 명시적으로 포착합니다. 또한 SceneBind Matching이라는 의미론적-공간적 매칭 방식을 제안합니다. 이 방식은 전역적인 장면 유사성과 객체 정렬을 통합하여 교차 모달 장면 검색 및 객체 기반 정보 추출을 지원합니다. SceneBind를 훈련하고 평가하기 위해, 구조화된 의미론적 및 공간적 주석이 포함된 새로운 실제 양이식 시각-청각 데이터 세트를 구축하고, 다양한 모달 간의 의미론적 및 공간적 신호를 정렬하는 훈련 프로토콜을 제안합니다. SceneBind는 대규모 사전 훈련된 의미론적 인코더와 호환되며, 몇 개의 추가 토큰만 사용하여 경량화된 공간 모델링을 제공합니다. 실험 결과, SceneBind는 최첨단 수준의 장면 및 공간 검색 성능을 달성했으며, 오디오-시각 로컬라이제이션과 같은 후속 작업에서 강력한 제로샷 전이 기능을 제공합니다.
We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spatial structure (i.e., where it is). SceneBind addresses this gap by representing each scene as a semantic-spatial entity, combining a global semantic embedding with object-centric semantic-spatial slots. This representation explicitly captures object-level semantics, spatial attributes, and uncertainty. We further propose SceneBind Matching, a semantic-spatial matching scheme that integrates global scene similarity with object alignment, supporting cross-modal scene retrieval and object grounding. To train and evaluate SceneBind, we curate a novel real-world binaural audio-visual dataset with structured semantic and spatial annotations, and propose a training protocol for aligning semantic and spatial signals across modalities. SceneBind is compatible with large-scale pretrained semantic encoders, adds lightweight spatial modeling with only a few additional tokens. It achieves state-of-the-art scene and spatial retrieval while enabling strong zero-shot transfer to downstream tasks such as audio-visual localization.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.