AffordTrajDP: 로봇 조작을 위한 동적 어포던스 기반 시각-운동 정책 학습
AffordTrajDP: Dynamic Affordance-Guided Visuomotor Policy Learning for Robotic Manipulation
어포던스 기반 모방 학습은 시각 정보를 작업별 기하학적 제약 조건(예: 고정된 접촉 지점)으로 압축하여 로봇 조작 작업에서 뛰어난 성능을 보여주었습니다. 그러나 일반적으로 사용되는 정적 어포던스는 정밀도가 중요한 작업이나 물체 위치 변화가 있는 경우 일관성을 잃어, 접촉 후 경로 편차를 초래할 수 있습니다. 이러한 문제를 해결하기 위해, 본 논문에서는 객체를 중심으로 시간적으로 정보를 전달하여 점진적인 조작 과정을 안내하는 동적 프레임워크인 AffordTrajDP를 제안합니다. 구체적으로, RGB-D 관찰 데이터를 기반으로, 저희는 핵심 아이디어로, 엔드 이펙터와 목표 물체 간의 원하는 접촉 지점을 나타내는 기준 어포던스를 활용하여, 물체의 SE(3) 자세를 자연스러운 정보 전달 매개체로 사용하여 시간적으로 일관되고 상태 정보를 고려한 안내를 제공하는 어포던스 경로를 생성합니다. AffordTrajDP는 ManiSkill3 데이터셋에서 평균 성공률 70.0%를 달성하며, 강력한 기준 모델보다 최대 17.8% 향상된 성능을 보였습니다. Galaxea A1 및 UR7e 로봇 팔을 사용한 실제 실험에서는 StackCube, PickCup, AdapterInsertion, Ring-on-Peg, Put-in-Bowl, USB Insertion 등 다양한 작업에서 물체 배치 변화 및 외관 변경에 대한 견고성을 검증했으며, Galaxea A1에서 학습된 데이터와 미지의 객체 인스턴스를 평가하고, 제안된 각 구성 요소의 기여도를 확인했습니다.
Affordance-guided imitation learning has shown impressive performance in robotic manipulation tasks by compressing visual perception into task-specific geometric constraints (e.g., fixed contact points). However, the commonly used static affordances can become inconsistent in precision-critical tasks or under object location perturbations, leading to post-contact trajectory drift. To address this issue, we propose AffordTrajDP, a dynamic framework that constructs affordance trajectories via object-centric temporal propagation to guide the progressive manipulation process. Specifically, given an RGB-D observation, our core insight is that a retrieved anchor affordance, which captures the desired contact point between the end-effector and the target object, can be propagated forward via affordance propagation, using the object's SE(3) pose as a natural propagation medium, to yield an affordance trajectory that provides temporally consistent, state-aware guidance throughout execution. AffordTrajDP achieves 70.0% average success rate on ManiSkill3, outperforming strong baselines by up to 17.8%. Real-world experiments on Galaxea A1 and UR7e robotic arms, covering StackCube, PickCup, AdapterInsertion, Ring-on-Peg, Put-in-Bowl, and USB Insertion, further validate robustness under object placement variations and appearance changes, with seen and unseen object instances evaluated on Galaxea A1, and ablations confirm the contribution of each proposed component.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.