객체를 오디오-비주얼 모달 음향장으로 표현
Objects as Audio-Visual Modal Sound Fields
현대적인 3차원 재구성 기술은 객체의 기하학적 구조와 외관을 모델링하는 데 뛰어나지만, 물리적인 상호작용을 통해 드러나는 풍부한 음향 정보를 대부분 무시합니다. 물체 충돌음은 재질, 강성 및 구조적 특성을 전달하며 시각 정보와 상호 보완되지만, 기존의 충돌음 모델링 방식은 고가의 물리 기반 시뮬레이션에 의존하거나 순수하게 데이터 중심적인 방식으로 일반화하기 위해서는 대규모 데이터셋이 필요합니다. 본 연구에서는 멀티 뷰 이미지와 소량의 충돌음 녹음을 사용하여 재구성된 새로운 객체 수준의 음향 표현인 Audio-Visual Modal Sound Field (AV-MSF)를 소개합니다. AV-MSF는 밀집된 3차원 시각적 특징과 통합된 3D 가우시안 스플래팅을 기반으로 강력한 기하학 정보를 제공하며, 압축되고 물리적으로 의미 있는 모달 파라미터를 사용하여 충돌음장을 표현함으로써 견고한 소량 데이터 재구성을 가능하게 합니다. 두 개의 실제 데이터셋에 대한 실험 결과, AV-MSF는 물리 기반 및 데이터 중심 방식의 기존 모델을 능가하는 최첨단 수준의 충돌음 렌더링 성능을 달성했습니다. 또한, 본 연구에서 제안하는 표현 방식을 활용하여 접촉 위치 파악 및 객체 음향 편집과 같은 다양한 응용 분야를 시연합니다.
While modern 3D reconstruction excels at modeling object geometry and appearance, it largely ignores the rich acoustic cues revealed through physical interaction. Object impact sounds convey material, stiffness, and structural properties that complement vision, yet existing impact sound modeling approaches either rely on expensive physics-based simulation or require large datasets to generalize in a purely data-driven manner. We introduce Audio-Visual Modal Sound Field (AV-MSF), a novel object-level acoustic representation reconstructed from multi-view images and only a few impact sound recordings. AV-MSF builds on 3D Gaussian Splatting integrated with dense 3D visual feature to provide a strong geometry-aware prior, and represents the impact sound field using compact, physically meaningful modal parameters, enabling robust few-shot reconstruction. Experiments on two real-world datasets show that AV-MSF achieves state-of-the-art impact sound rendering, outperforming both physics-based and data-driven baselines. Furthermore, we demonstrate downstream applications enabled by our representation, including contact localization and object sound editing.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.