2606.31167v1 Jun 30, 2026 cs.RO

MIRTH: 상호 정보 추론을 활용한 시간적 허브 기반 시각-언어-행동 에이전트

MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents

Ziwei Niu
Ziwei Niu
Citations: 198
h-index: 8
Shiyu Teng
Shiyu Teng
Citations: 291
h-index: 7
Hao Sun
Hao Sun
Citations: 36
h-index: 2
Yu Song
Yu Song
Citations: 34
h-index: 3
Yen-wei Chen
Yen-wei Chen
Citations: 304
h-index: 10

VLA(Vision-Language-Action) 모델은 웹 규모의 데이터에서 얻은 의미 정보를 물리적인 로봇 제어로 전달하는 강력한 패러다임으로 부상했습니다. 그러나 현재 단일 프레임 기반 아키텍처는 다음과 같은 근본적인 한계점을 가지고 있습니다: 과거 동역학을 무시하는 시간적 인식 부족, 고수준 명령과 저수준 모터 명령 간의 추론 격차, 그리고 오토리그레시브 스칼라 디코딩으로 인한 비효율적인 추론. 본 연구에서는 이러한 과제를 해결하기 위해 설계된 통합 프레임워크인 MIRTH를 제안합니다. MIRTH는 사전 훈련된 VLA 백본에 다음 세 가지 핵심 혁신을 추가합니다: (1) 장기적인 장면 변화와 단기적인 동작 경향을 압축하여 간결한 임베딩으로 표현하는 이중 스케일 시간적 메모리 허브; (2) 상호 정보 기반 최적화를 통해 생성된 잠재 추론 토큰은 다중 모달 컨텍스트를 행동 궤적으로 정렬하기 위한 의미 계획 공간을 구축합니다; (3) 오토리그레시브 생성을 벡터 단위 예측으로 대체하여 제어 처리량을 극대화하는 병렬 액션 디코딩 방식. LIBERO 시뮬레이션 벤치마크 및 실제 LeRobot 플랫폼에서의 광범위한 실험 결과, MIRTH는 최첨단 성능을 달성하며, 예상치 못한 오류 복구 능력을 보여주었습니다. 코드와 수집된 데이터셋은 http://github.com/kiva12138/mirth 에서 공개됩니다.

Original Abstract

VLA models have emerged as a powerful paradigm for transferring semantic knowledge from web-scale data to physical robotic control. However, current single-frame architectures suffer from intrinsic limitations: temporal myopia that discards historical dynamics, reasoning gaps between high-level instructions and low-level motor commands, and inference inefficiency due to autoregressive scalar decoding. In this work, we propose MIRTH, a unified framework designed to address these challenges. MIRTH augments a pretrained VLA backbone with three key innovations: (1) dual-scale temporal memory hubs that compress long-term scene evolution and short-term motion trends into compact embeddings; (2) latent reasoning tokens optimized via a mutual-information objective carving out a semantic plan space to align multimodal context with action trajectories; and (3) a parallel action decoding scheme that replaces autoregressive generation with vector-wise prediction to maximize control throughput. Extensive evaluations on the LIBERO simulation benchmark and a real-world LeRobot platform demonstrate that MIRTH achieves state-of-the-art performance and exhibiting emergent error recovery capabilities. The codes and collected datasets are released at http://github.com/kiva12138/mirth.

0 Citations
0 Influential
28.4657359028 Altmetric
0.0 Score
Original PDF
1

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!