TouchThinker: 대규모 데이터와 동작 인지 표현을 활용하여 촉각 기반 상식 추론을 개방형 환경으로 확장
TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation
촉각은 로봇 에이전트가 물리적 세계를 이해하는 데 중요한 역할을 합니다. 최근 연구에서는 촉각 신호를 언어 시스템에 통합하여 촉각 기반 상식 추론을 수행하려는 시도가 있었지만, 실제 개방형 환경으로 이러한 시스템을 확장하는 것은 두 가지 주요 문제점으로 인해 여전히 어렵습니다. (1) 현재의 촉각 추론 데이터셋은 형식과 규모 면에서 제한적이며, 이는 촉각 관찰로부터 물리적 상식을 추론하는 데 필요한 충분한 학습 데이터를 제공하지 못하고, 일반화 가능한 촉각 상식 학습을 방해합니다. (2) 촉각 신호는 본질적으로 중복적이고 행동에 특이적이기는 하지만, 기존 방법에서는 이러한 특징을 간과하여 비효율적인 표현 방식을 사용하며, 이는 제한적인 의미적 표현력을 갖게 됩니다. 이러한 한계점을 극복하기 위해, 우리는 데이터와 표현 방식 모두에서 촉각 기반 상식 추론을 개방형 환경으로 확장하는 촉각-언어 프레임워크인 TouchThinker를 제안합니다. 첫째, 우리는 415개의 객체, 8가지 시나리오, 그리고 7가지 센서 유형을 포함하는 백만 규모의 다중 소스 촉각 추론 데이터셋인 TouchThinker-1M을 구축하여 개방형 환경에서의 일반화에 필요한 견고한 데이터 기반을 제공합니다. 또한, 더욱 현실적이고 다양한 작업을 포함하는 개방형 벤치마크인 TouchThinker-Bench를 소개합니다. 둘째, 우리는 촉각 표현의 효율성을 향상시키고 효율적인 추론을 가능하게 하는 동작 인지 모델링 메커니즘을 제안합니다. 실험 결과는 TouchThinker가 여러 데이터셋에서 최첨단 모델과 경쟁력 있는 성능을 달성함을 보여줍니다. 저희의 코드와 데이터셋은 다음 주소에서 제공됩니다: https://github.com/lvkailin0118/TouchThinker.
Touch is a key modality for embodied agents to understand the physical world. Although recent work has incorporated tactile signals into language systems for tactile commonsense reasoning, scaling such systems to realistic open-world settings remains challenging due to two key bottlenecks: (1) current tactile reasoning datasets remain limited in format and scale, providing insufficient supervision for reasoning from tactile observations to physical commonsense and hindering the learning of transferable tactile commonsense; (2) Tactile signals are inherently redundant and action-specific, yet existing methods often overlook these properties, resulting in inefficient representations with limited semantic expressiveness. To address these limitations, we propose TouchThinker, a tactile-language framework that scales tactile commonsense reasoning to the open world from both data and representation perspectives. First, we construct TouchThinker-1M, a million-scale, multi-source tactile reasoning dataset covering \textbf{415} objects, \textbf{8} scenarios, and \textbf{7} sensor types, providing a solid data foundation for open-world generalization. We further introduce TouchThinker-Bench, an open-world benchmark with more realistic and diverse tasks. Then, we propose action-aware modeling mechanism to improve tactile representation efficiency and enable efficient reasoning. Experimental results demonstrate that TouchThinker achieves competitive performance against state-of-the-art models across multiple datasets. Our code and dataset will be made available at: https://github.com/lvkailin0118/TouchThinker.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.