2607.05780v1 Jul 07, 2026 cs.RO

FORGE: 핵심 지점 궤적 추론을 통한 기능 도구 사용의 일반화

FORGE: Towards Functional Tool-Use Generalization via Keypoint Trajectory Reasoning

Xiangyu Chen
Xiangyu Chen
Citations: 9
h-index: 1
Yuxuan Hu
Yuxuan Hu
Citations: 11
h-index: 2
Jianfei Yang
Jianfei Yang
Citations: 27
h-index: 2
Shuxin Cao
Shuxin Cao
Citations: 120
h-index: 4
Chuhao Zhou
Chuhao Zhou
Citations: 19
h-index: 2
Liquang Wang
Liquang Wang
Citations: 162
h-index: 5
Boyu Ma
Boyu Ma
Citations: 89
h-index: 4
Animesh Garg
Animesh Garg
Citations: 377
h-index: 7

사람은 책, 돌 또는 신발과 같은 물건을 활용하여 못을 박는 등 다양한 용도로 재활용할 수 있지만, 특정 도구를 학습한 로봇은 동일한 기능을 새로운 도구에 적용하는 데 어려움을 겪습니다. 이는 기능적 일반화의 격차를 나타냅니다. 이러한 도구들은 시각적으로 인식 가능한 공통된 기능적 의도를 공유하지만, 이 인지적 유사성은 행동 공간으로 전달되지 않으며, 각 도구는 완전히 다른 운동 패턴을 요구합니다. 이러한 격차를 해소하기 위해 우리는 어포던스 이미지, 인간 비디오 프롬프트 및 2D 핵심 지점 궤적과 같은 중간 표현을 탐색했으며, 핵심 지점 궤적이 기능적 표현력과 행동 기반의 안정성을 가장 잘 균형을 이루는 것을 확인했습니다. 이를 바탕으로, 저희는 기능적 추론과 실행을 분리하는 두 단계 정책인 FunctiOnal Reasoning and Grounded Execution (FORGE)를 제안합니다. FORGE는 동작 정보가 없는 데이터로부터 일반화 가능한 핵심 지점 궤적을 예측한 다음, 제한된 시연 데이터를 기반으로 이를 로봇의 행동으로 연결합니다. 저희는 일곱 가지 도구를 사용한 타격 기능 벤치마크에서, FORGE가 최첨단 방법보다 실제 및 시뮬레이션 환경 모두에서 새로운 도구에 대해 지속적으로 우수한 성능을 보이며, 평균 성공률이 2배 이상 향상되었습니다.

Original Abstract

While humans readily repurpose a book, a stone, or a shoe to drive a nail, robots trained on specific tools fail to transfer the same function to novel ones -- a gap we formalize as functional generalization. Such tools share a common functional intent that is visually recognizable, yet this perceptual similarity does not carry over to action space, where each tool demands an entirely different motor pattern. To bridge this gap, we explore intermediate representations including affordance images, human video prompts, and 2D keypoint trajectories, finding that keypoint trajectories best balance functional expressiveness and action groundability. Building on this, we propose FunctiOnal Reasoning and Grounded Execution (FORGE), a two-stage policy that decouples functional reasoning from action execution: predicting generalizable keypoint trajectories from action-free data, then grounding them into robot actions with limited demonstrations. On a seven-tool hitting-function benchmark, FORGE consistently outperforms state-of-the-art methods on unseen tools in both simulation and the real world, achieving over 2X improvement in average success rate.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!