TAM: 조작을 위한 견고한 동작 전송을 위한 토크 적응 모듈
TAM: Torque Adaptation Module for Robust Motion Transfer in Manipulation
하나의 로봇에 대해 튜닝된 정책은 시뮬레이션과 실제 환경 간의 차이, 알려지지 않은 하중 또는 동일한 로봇 인스턴스의 다른 동역학으로 인해 다른 로봇에서 다르게 작동할 수 있습니다. 접촉이 많은 동적 조작에서는 작은 동작 불일치도 레퍼런스 동작 추적 실패로 이어질 수 있는데, 이는 접촉의 타이밍과 방식을 방해하기 때문입니다. 일반적인 해결책인 도메인 랜덤화 또는 시스템 식별은 과도하게 보수적인 작업 정책을 생성하거나 각 로봇 또는 하중에 대해 다시 수집해야 하는 데이터를 필요로 합니다. 우리는 토크 적응 모듈(TAM)을 소개합니다. TAM은 로봇에 전달되는 토크 명령을 조정하여 이상적인 로봇의 동작과 일치하도록 학습된 모듈입니다. TAM은 정책의 동작을 추적하는 저수준 제어기와 로봇의 토크 인터페이스 사이에 작동하며, 고유한 정보를 담고 있는 과거 기록을 잠재 상태로 포함하는 히스토리 인코더와 잔여 토크 보정을 계산하는 토크 어댑터를 포함합니다. TAM은 정책 관찰이나 동작 공간에 의존하지 않고 고유 정보의 과거 기록에만 의존하기 때문에 동일한 TAM 가중치를 사용하여 서로 다른 동작 공간(조인트 목표, 엔드 이펙터 목표 또는 직접 토크)을 갖는 정책을 조정할 수 있습니다. 정책 자체는 로봇 파라미터의 도메인 랜덤화로 훈련할 필요가 없습니다. 대신, 우리는 로봇 파라미터의 도메인 랜덤화를 TAM에 위임하고, 멀티 로봇 사전 훈련 후 로봇별 미세 조정을 통해 완전히 랜덤화된 시뮬레이션에서 학습시킵니다. 이 단계는 실제 로봇 데이터를 전혀 사용하지 않습니다. 우리는 TAM을 Franka Panda 로봇에서 다양한 동적 조작 작업(RL 기반 비전 박스 밀기 정책, BC 기반 플립 정책 및 MPC 볼-온-플레이트 균형)에 대해 0샷으로 평가했습니다. 실험 결과, TAM은 온라인 시스템 식별 및 RMA 기준보다 향상된 0샷 실제 로봇 실행 성능을 제공하며 견고한 동적 조작 성능을 가능하게 합니다.
A policy tuned for one robot often behaves differently on another, whether due to the sim-to-real gap, unknown payloads, or the differing dynamics of two instances of the same robot. In contact-rich, dynamic manipulation, even small motion discrepancies can result in failure to track reference motion, since they disrupt the timing and modes of contact. Common remedies, such as domain randomization or system identification, either produce overly conservative task policies or require data that must be recollected for each robot or payload. We introduce the Torque Adaptation Module (TAM), a learned module that adapts the torque commands sent to the robot to match the behavior of an ideal robot. TAM operates between the low-level controller that tracks the policy's actions and the robot's torque interface. It includes a history encoder that embeds proprioceptive history into a latent state and a torque adaptor that computes residual torque corrections. Because TAM depends only on proprioceptive history and not on policy observations, or the action space, the same TAM weights can be reused to adapt policies with different action spaces (joint targets, end-effector targets, or direct torques). The policies themselves do not need to be trained with domain randomization of robot parameters. Instead, we offload the need for domain randomization to TAM by training it entirely in randomized simulation, using multi-robot pretraining followed by a robot-specific fine-tuning step that still requires no real-robot data. We evaluate TAM zero-shot on a real Franka Panda robot across dynamic manipulation tasks that include a vision-based box pushing policy (from RL), a flip policy (from BC), and an MPC ball-on-plate balancing. Our experiments show that TAM improves zero-shot real-robot execution compared to online system identification and RMA baselines and enables robust dynamic manipulation performance.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.