2607.15265v1 Jul 16, 2026 cs.CV

SceneBind: 비전, 오디오 및 언어를 통해 무엇과 어디를 연결하는 방법

SceneBind: Binding What and Where Across Vision, Audio and Language

Mingfei Chen
Mingfei Chen
Citations: 155
h-index: 6
Zijun Cui
Zijun Cui
Citations: 30
h-index: 2
Ruoke Zhang
Ruoke Zhang
Citations: 0
h-index: 0
H. Ryu
H. Ryu
Citations: 241
h-index: 7
Eli Shlizerman
Eli Shlizerman
Citations: 2,259
h-index: 23

본 논문에서는 SceneBind를 제안합니다. SceneBind는 현실적인 장면을 전반적인 의미론적 정보와 3차원 공간 정보를 결합하여 표현하며, 이는 시각, 음성 및 언어 데이터를 포괄합니다. 기존의 다중 모달 인코더는 개별 객체의 의미론적 정보(즉, 무엇이 존재하는가)를 잘 파악하지만, 종종 명시적인 공간 구조(즉, 어디에 있는가)에 대한 이해가 부족합니다. SceneBind는 이러한 간극을 메우기 위해 각 장면을 전역 의미론적 임베딩과 객체 중심의 의미론적-공간적 슬롯을 결합한 의미론적-공간적 개체로 표현합니다. 이 표현은 객체 수준의 의미, 공간 속성 및 불확실성을 명시적으로 포착합니다. 또한 SceneBind Matching이라는 의미론적-공간적 매칭 방식을 제안합니다. 이 방식은 전역적인 장면 유사성과 객체 정렬을 통합하여 교차 모달 장면 검색 및 객체 기반 정보 추출을 지원합니다. SceneBind를 훈련하고 평가하기 위해, 구조화된 의미론적 및 공간적 주석이 포함된 새로운 실제 양이식 시각-청각 데이터 세트를 구축하고, 다양한 모달 간의 의미론적 및 공간적 신호를 정렬하는 훈련 프로토콜을 제안합니다. SceneBind는 대규모 사전 훈련된 의미론적 인코더와 호환되며, 몇 개의 추가 토큰만 사용하여 경량화된 공간 모델링을 제공합니다. 실험 결과, SceneBind는 최첨단 수준의 장면 및 공간 검색 성능을 달성했으며, 오디오-시각 로컬라이제이션과 같은 후속 작업에서 강력한 제로샷 전이 기능을 제공합니다.

Original Abstract

We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spatial structure (i.e., where it is). SceneBind addresses this gap by representing each scene as a semantic-spatial entity, combining a global semantic embedding with object-centric semantic-spatial slots. This representation explicitly captures object-level semantics, spatial attributes, and uncertainty. We further propose SceneBind Matching, a semantic-spatial matching scheme that integrates global scene similarity with object alignment, supporting cross-modal scene retrieval and object grounding. To train and evaluate SceneBind, we curate a novel real-world binaural audio-visual dataset with structured semantic and spatial annotations, and propose a training protocol for aligning semantic and spatial signals across modalities. SceneBind is compatible with large-scale pretrained semantic encoders, adds lightweight spatial modeling with only a few additional tokens. It achieves state-of-the-art scene and spatial retrieval while enabling strong zero-shot transfer to downstream tasks such as audio-visual localization.

0 Citations
0 Influential
11.5 Altmetric
57.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!