2603.09882v1 Mar 10, 2026 cs.RO

혼잡한 환경에서 동역학 기반 정책 학습을 통한 새로운 유형의 외부 조작 능력

Emerging Extrinsic Dexterity in Cluttered Scenes via Dynamics-aware Policy Learning

Yixin Zheng
Yixin Zheng
Citations: 44
h-index: 5
Jiangran Lyu
Jiangran Lyu
Citations: 184
h-index: 7
Yu Deng
Yu Deng
Citations: 4
h-index: 1
Xuesong Shi
Xuesong Shi
Citations: 45
h-index: 4
Xiaoguang Zhao
Xiaoguang Zhao
Citations: 19
h-index: 3
Yizhou Wang
Yizhou Wang
Citations: 1
h-index: 1
Zhizheng Zhang
Zhizheng Zhang
Citations: 958
h-index: 10
Yifan Zhang
Yifan Zhang
Citations: 28
h-index: 3
Jiayi Chen
Jiayi Chen
Citations: 291
h-index: 7
Mi Yan
Mi Yan
Citations: 128
h-index: 3
He Wang
He Wang
Citations: 27
h-index: 3

외부 조작 능력은 환경과의 접촉을 활용하여 기존의 잡기 조작의 한계를 극복합니다. 그러나 여러 상호 작용하는 물체들의 복잡하게 결합된 동역학을 선택적으로 활용해야 하기 때문에, 혼잡한 환경에서 이러한 조작 능력을 달성하는 것은 여전히 어려운 과제이며, 이에 대한 연구가 부족합니다. 기존의 접근 방식은 이러한 복잡한 동역학을 명시적으로 모델링하지 못하기 때문에, 혼잡한 환경에서의 비-잡기 조작에 한계가 있으며, 이는 실제 환경에서의 활용 가능성을 제한합니다. 본 논문에서는 혼잡한 환경에서 접촉에 의해 발생하는 물체의 동역학을 학습된 표현으로 활용하여 정책 학습을 용이하게 하는 동역학 기반 정책 학습 (Dynamics-Aware Policy Learning, DAPL) 프레임워크를 소개합니다. 이 표현은 명시적인 세계 모델링을 통해 학습되며, 강화 학습을 위한 조건으로 사용되어, 사람이 설계한 접촉 규칙이나 복잡한 보상 설계 없이 외부 조작 능력이 발현되도록 합니다. 저희는 이 방법을 시뮬레이션 환경과 실제 환경 모두에서 평가했습니다. 저희 방법은 다양한 밀도의 새로운 시뮬레이션된 혼잡한 환경에서 성공률이 25% 이상 높으며, 기존의 잡기 조작, 인간 원격 조작, 그리고 기존의 표현 기반 정책보다 우수한 성능을 보였습니다. 실제 환경에서는 10개의 혼잡한 환경에서 약 50%의 성공률을 달성했으며, 실제 식료품점에서 활용하는 실험을 통해 강력한 시뮬레이션-실제 환경 간의 전이와 적용 가능성을 입증했습니다.

Original Abstract

Extrinsic dexterity leverages environmental contact to overcome the limitations of prehensile manipulation. However, achieving such dexterity in cluttered scenes remains challenging and underexplored, as it requires selectively exploiting contact among multiple interacting objects with inherently coupled dynamics. Existing approaches lack explicit modeling of such complex dynamics and therefore fall short in non-prehensile manipulation in cluttered environments, which in turn limits their practical applicability in real-world environments. In this paper, we introduce a Dynamics-Aware Policy Learning (DAPL) framework that can facilitate policy learning with a learned representation of contact-induced object dynamics in cluttered environments. This representation is learned through explicit world modeling and used to condition reinforcement learning, enabling extrinsic dexterity to emerge without hand-crafted contact heuristics or complex reward shaping. We evaluate our approach in both simulation and the real world. Our method outperforms prehensile manipulation, human teleoperation, and prior representation-based policies by over 25% in success rate on unseen simulated cluttered scenes with varying densities. The real-world success rate reaches around 50% across 10 cluttered scenes, while a practical grocery deployment further demonstrates robust sim-to-real transfer and applicability.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!