2602.13640v1 Feb 14, 2026 cs.RO

정밀 로봇 조작을 위한 계층적 오디오-시각-고유수용 융합

Hierarchical Audio-Visual-Proprioceptive Fusion for Precise Robotic Manipulation

Siyuan Li
Siyuan Li
Citations: 67
h-index: 4
Jiani Lu
Jiani Lu
Citations: 19
h-index: 2
Yu Song
Yu Song
Citations: 0
h-index: 0
Xianren Li
Xianren Li
Citations: 0
h-index: 0
Bo An
Bo An
Citations: 2
h-index: 1
Peng Liu
Peng Liu
Citations: 47
h-index: 4

기존의 로봇 조작 방법은 주로 시각 정보와 고유수용성 정보를 기반으로 하지만, 부분적으로 관찰 가능한 실제 환경에서는 접촉과 관련된 상호 작용 상태를 추론하는 데 어려움이 있을 수 있습니다. 반면, 음향 정보는 접촉 과정에서 풍부한 상호 작용 역학을 자연스럽게 포함하지만, 현재의 다중 모드 융합 연구에서는 아직 충분히 활용되지 않고 있습니다. 대부분의 다중 모드 융합 방법은 모드 간의 역할을 동일하다고 가정하고, 평탄하고 대칭적인 융합 구조를 설계합니다. 그러나 이러한 가정은 본질적으로 희소하고 접촉에 의존적인 음향 신호에는 적합하지 않습니다. 본 연구에서는 음향 정보를 활용한 정밀 로봇 조작을 위해, 오디오, 시각, 그리고 고유수용성을 점진적으로 통합하는 계층적 표현 융합 프레임워크를 제안합니다. 제안하는 방법은 먼저 음향 정보를 사용하여 시각 및 고유수용성 표현을 조건부로 생성하고, 이후 고차의 모드 간 상호 작용을 명시적으로 모델링하여 모드 간의 상호 보완적인 의존성을 파악합니다. 융합된 표현은 확산 기반 정책에 의해 활용되어, 다중 모드 관찰로부터 연속적인 로봇 동작을 직접 생성합니다. 엔드투엔드 학습과 계층적 융합 구조의 결합은 정책이 작업과 관련된 음향 정보를 활용하고, 상대적으로 정보량이 적은 모드로부터의 간섭을 완화할 수 있도록 합니다. 제안하는 방법은 실제 로봇 조작 작업, 예를 들어 액체 따르기 및 캐비닛 열기 작업에 대해 평가되었습니다. 광범위한 실험 결과는 제안하는 방법이 최첨단 다중 모드 융합 프레임워크보다 일관되게 우수한 성능을 보이며, 특히 시각 정보만으로는 쉽게 얻을 수 없는 작업 관련 정보를 음향 정보가 제공하는 시나리오에서 더욱 두드러진 성능 향상을 보임을 보여줍니다. 또한, 다중 모드 융합을 통한 로봇 조작에서 음향 정보의 효과를 해석하기 위해 상호 정보 분석을 수행했습니다.

Original Abstract

Existing robotic manipulation methods primarily rely on visual and proprioceptive observations, which may struggle to infer contact-related interaction states in partially observable real-world environments. Acoustic cues, by contrast, naturally encode rich interaction dynamics during contact, yet remain underexploited in current multimodal fusion literature. Most multimodal fusion approaches implicitly assume homogeneous roles across modalities, and thus design flat and symmetric fusion structures. However, this assumption is ill-suited for acoustic signals, which are inherently sparse and contact-driven. To achieve precise robotic manipulation through acoustic-informed perception, we propose a hierarchical representation fusion framework that progressively integrates audio, vision, and proprioception. Our approach first conditions visual and proprioceptive representations on acoustic cues, and then explicitly models higher-order cross-modal interactions to capture complementary dependencies among modalities. The fused representation is leveraged by a diffusion-based policy to directly generate continuous robot actions from multimodal observations. The combination of end-to-end learning and hierarchical fusion structure enables the policy to exploit task-relevant acoustic information while mitigating interference from less informative modalities. The proposed method has been evaluated on real-world robotic manipulation tasks, including liquid pouring and cabinet opening. Extensive experiment results demonstrate that our approach consistently outperforms state-of-the-art multimodal fusion frameworks, particularly in scenarios where acoustic cues provide task-relevant information not readily available from visual observations alone. Furthermore, a mutual information analysis is conducted to interpret the effect of audio cues in robotic manipulation via multimodal fusion.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!