FA-RDP: 접촉이 빈번한 조작을 위한 주파수 적응형 반응 확산 정책
FA-RDP: A Frequency-Adaptive Reactive Diffusion Policy for Contact-Rich Manipulation
접촉이 빈번한 조작 환경에서, 다양한 액션 모드가 존재하며, 이는 하나의 에피소드의 여러 단계에서 중요한 역할을 합니다. 접촉 전에는 여러 경로가 동등하게 유효할 수 있으므로, 다양한 액션 방식을 유지하는 것이 중요합니다. 반면, 접촉 후에는 기하학적 제약 조건과 힘 제한으로 인해 해 공간이 좁아지며, 성공적인 수행을 위해서는 힘 피드백에 대한 빠른 반응이 필요합니다. 그러나 기존의 확산 정책은 에피소드 전체에 걸쳐 고정된 추론 주파수와 샘플링 단계를 사용하므로, 근본적인 타협점을 갖습니다. 낮은 주파수의 다단계 샘플링은 접촉 전의 다양한 액션 방식을 더 잘 보존하지만, 힘 피드백에 대한 반응 속도가 느린 반면, 높은 주파수의 샘플링은 반응성을 향상시키지만, 접촉 전에 존재하는 뚜렷한 액션 방식을 소실시키는 경향이 있습니다. 이러한 상충 관계를 해결하기 위해, 우리는 주파수 적응형 반응 확산 정책인 FA-RDP를 제안합니다. 공유된 다중 주파수 시각 및 힘 변환기는 낮은 주파수와 높은 주파수 모두에서 액션 덩어리를 예측하며, 학습된 다모드 지표는 접촉 전에 액션의 불확실성이 감소함에 따라 다단계 저주파 샘플링과 단일 단계 고주파 샘플링을 동적으로 선택합니다. 또한, 우리는 확산 네트워크를 재구성하여 로봇 액션 공간에서 액션을 예측하면서 DDPM 기반 잔차 감독을 유지하는 Manifold Consistency Distillation (MCD) 방법을 도입했습니다. 세 가지 접촉이 빈번한 조작 작업에 대한 실험 결과는 FA-RDP가 가장 높은 성공률을 달성했으며, 동시에 접촉 전의 다양한 경로 방식을 보존함을 보여줍니다. 코드와 동영상은 https://fa-rdp.github.io 에서 확인할 수 있습니다.
In contact-rich manipulation, action multimodality and reactivity dominate different stages of a single episode. Before contact, multiple trajectories might be equally valid, making it important to preserve diverse action modes. After contact, geometric constraints and force limits narrow the solution space, while successful execution demands rapid responses to force feedback. However, standard diffusion policies use a fixed inference frequency and sampling steps throughout the episode, forcing a fundamental compromise: low-frequency, multi-step sampling better preserves pre-contact multimodality but responds slowly to force feedback, whereas high-frequency sampling improves reactivity but tends to collapse distinct pre-contact modes. To resolve this tradeoff, we present FA-RDP, a frequency-adaptive reactive diffusion policy. A shared multi-frequency visual-force Transformer predicts action chunks at both low and high frequencies, while a learned multimodality indicator dynamically selects multi-step low-frequency sampling before contact and one-step high-frequency sampling as action ambiguity decreases. We further introduce Manifold Consistency Distillation (MCD), which reparameterizes the diffusion network to predict actions on the robot action manifold while retaining DDPM-based residual supervision. Experiments on three contact-rich manipulation tasks show that FA-RDP achieves the highest success rate while preserving diverse pre-contact trajectory modes. Code and videos are available at https://fa-rdp.github.io.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.