2607.26460v1 Jul 29, 2026 cs.RO

RLMM-Flow: 잠재 공간 강화 학습을 활용한 흐름 기반 모바일 매니퓰레이션 프레임워크

RLMM-Flow: A Flow-based Mobile Manipulation Framework with Latent-Space Reinforcement Learning

Hui Cheng
Hui Cheng
Citations: 79
h-index: 4
Shuhang Wang
Shuhang Wang
Citations: 26
h-index: 2
Ziming Li
Ziming Li
Citations: 0
h-index: 0

모바일 매니퓰레이션은 목표 달성, 충돌 회피, 베이스의 운동학적 제약 조건, 조인트 제한 및 궤적 부드러움을 동시에 만족하는 전체 동작 시퀀스를 생성해야 합니다. 흐름 기반 생성 정책은 전문가 데모로부터 다양한 모드와 시간적으로 일관된 동작 패턴을 학습하는 효율적인 방법을 제공하지만, 단순히 모방 학습만으로는 정책의 품질을 향상시킬 수 없습니다. 본 연구에서는 전문가 흐름 정책의 사전 훈련과 잠재 공간 강화 학습 후 훈련을 결합한 흐름 기반 모바일 매니퓰레이션 프레임워크인 RLMM-Flow를 제안합니다. 이 프레임워크는 먼저 전문가 데모로부터 다양한 전체 동작 패턴을 학습하는 흐름 정책을 학습합니다. 사전 훈련된 흐름 정책은 고정된 상태로 유지되며, 잠재 공간 조향 네트워크가 초기 노이즈를 더 높은 가치를 갖는 동작 시퀀스로 조종합니다. 고차원 잠재 최적화를 안정화하기 위해, 잠재 공간 액터와 크리틱을 공동 훈련하기 전에 액션 공간에 대한 크리틱을 먼저 학습시키고, 점진적으로 제어를 확대하는 방식의 세분화된 잠재 공간 조향 방식을 도입하여, 초기에는 전체 시간 범위에서 공유되는 잠재 표현부터 시작하여 최종적으로는 완전한 차원의 잔차 표현까지 확장합니다. 모바일 매니퓰레이션 동작 계획 벤치마크 실험 결과, RLMM-Flow는 단순 모방 학습 기반의 흐름 정책 및 기존 강화 학습 후 훈련 방법과 비교했을 때 작업 성공률, 충돌 회피 능력 및 궤적 품질을 크게 향상시키며, 동시에 빠른 흐름 기반 추론 속도를 유지합니다.

Original Abstract

Mobile manipulation requires generating whole-body action chunks that jointly satisfy goal reaching, collision avoidance, base kinematic constraints, manipulator joint limits, and trajectory smoothness. Flow-based generative policies provide an efficient paradigm for learning multimodal and temporally consistent motion priors from expert demonstrations, but imitation-only training cannot improve policy quality beyond the demonstration distribution. We propose RLMM-Flow, a flow-based mobile manipulation framework that combines expert flow-policy pretraining with latent-space reinforcement learning post-training. The framework first learns a flow policy that captures a multimodal whole-body motion prior from expert demonstrations. The pretrained flow policy is then frozen, while a latent steering network steers its initial noise toward higher-value action chunks. To stabilize high-dimensional latent optimization, we warm up an action-space critic before jointly training the latent critic and latent actor, and introduce coarse-to-fine latent steering that progressively expands control from a horizon-shared latent representation to a full-dimensional residual representation. Experiments on mobile manipulation motion-planning benchmarks show that RLMM-Flow substantially improves task success, collision avoidance, and trajectory quality over imitation-only flow policies and existing reinforcement learning post-training baselines, while preserving fast flow-based inference.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!