TacReasoner: 실제 환경에서의 상호작용적 추론을 위한 동적 촉각 언어 프레임워크
TacReasoner: A Dynamic Tactile-Language Framework for Interactive Reasoning in Real-World Scenarios
오감 중 촉각은 물리적인 접촉과 상호작용을 인식하는 데 필수적이므로 생존에 가장 중요한 감각이라고 할 수 있습니다. 본 논문에서는 다중 모드 추론을 위한 지능형 시스템에 촉각 센싱을 통합하는 데 있어 두 가지 주요 과제를 탐구합니다. (i) 동적 촉각 신호의 불충분한 모델링은 시간적으로 변화하는 속성에 대한 추론을 제한하며, (ii) 명시적인 추론 메커니즘이 없기 때문에 촉각 기반 모델에서 환각 현상이 발생하여 실제 환경에서의 추론 안정성이 저하됩니다. 이러한 과제를 해결하기 위해, 본 논문에서는 실제 환경에서의 상호작용적 추론을 위한 동적 촉각 언어 프레임워크인 TacReasoner를 제안합니다. 먼저, TacReasoner는 동적 속도에 대한 인지 및 표현을 향상시키기 위해 Dynamic-aware Tactile Encoder를 통합합니다. 더욱 중요하게는, 우리는 구조화된 촉각 입력에 대한 체계적인 추론을 위한 최초의 촉각 체인 오브 씽크(Chain-of-Thought) 데이터셋인 TouchCoT-10k를 소개합니다. 이를 기반으로, 우리는 동적 촉각 인지와 실제 환경 상식 추론을 체계적으로 평가하기 위해 DynTac-Bench를 구축했습니다. 실험 결과는 TacReasoner가 여러 데이터셋에서 최첨단 모델과 경쟁력 있는 성능을 달성함을 보여줍니다. 특히, 70억 개의 파라미터만 사용했음에도 불구하고 TacReasoner는 대부분의 하위 작업에서 140억 개의 파라미터를 가진 VTV-LLM 모델보다 뛰어난 성능을 보이며, 촉각 상식 추론에서의 효과성과 효율성을 강조합니다.
Among the five primary human senses, tactile is arguably the most fundamental to survival, as it enables the perception of physical contact and interaction in real-world environments. In this paper, we explore two key challenges of integrating tactile sensing into intelligent systems for multimodal reasoning: (i) insufficient modeling of dynamic tactile signals, which restricts reasoning over temporally evolving properties, and (ii) hallucination in tactile foundation models caused by the absence of explicit reasoning mechanisms, leading to unstable real-world inference. To address these challenges, we propose TacReasoner, a dynamic tactile-language framework for interactive reasoning in real-world scenarios. First, TacReasoner incorporates a Dynamic-aware Tactile Encoder to enhance the perception and representation of dynamic tactile signals. More importantly, we introduce TouchCoT-10k, the first tactile chain-of-thought dataset for structured reasoning over tactile inputs. Upon it, we establish DynTac-Bench to systematically evaluate dynamic tactile perception and real-world commonsense reasoning. Experimental results demonstrate that TacReasoner achieves competitive performance against state-of-the-art models across multiple datasets. Notably, despite using only 7B parameters, TacReasoner outperforms the 14B VTV-LLM model on most subtasks, highlighting its effectiveness and efficiency in tactile commonsense reasoning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.