FloAff-Kitchen: 표준화된 및 점진적인 바닥 활용도 학습을 통한 탐색 및 조작 간의 연결
FloAff-Kitchen: Bridging Navigation and Manipulation via Canonical and Progressive Floor Affordance Learning
로봇이 단순히 이동 가능성을 보장하는 것 이상으로, 후속 조작 성공률을 극대화하는 바닥 활용도(FloAff)를 식별해야 하는 것이 모바일 조작의 핵심입니다. FloAff 예측은 목표 지향적이고 지역적인 공간 추론 문제이지만, 기존 방법들은 관련 없는 공간 정보와 임의적인 객체 방향으로 인한 표현의 모호성 문제를 겪고 있으며, 다양한 조작 기술에 걸쳐 공유되는 지식과 작업별 특정 지식을 분리하는 데 어려움을 겪습니다. 이러한 과제를 해결하기 위해, 우리는 중심 시점에서 얻은 다중 감각 정보를 기반으로 FloAff 예측을 위한 통합 프레임워크를 제안합니다. 이 프레임워크는 표준화된 표현 학습과 점진적인 활용도 사전 학습으로 구성됩니다. 구체적으로, 우리는 로봇의 베이스 위치와 관련된 중요한 지역 구조는 유지하면서 관련 없는 공간 변화를 제거하여 표준적인 상호 작용 형상을 학습하는 '표준 바닥 활용도 표현(CFAR)'을 도입했습니다. 또한, 우리는 기초 조작 작업에서 전이 가능한 FloAff 사전 지식을 학습하고, 이를 다양한 후속 조작 기술에 점진적으로 적용하는 '점진적 바닥 활용도 학습(PFAL)'을 제안합니다. 체계적인 평가를 위해, 우리는 다양한 조작 기술, 장면 구성, 가구 스타일 및 시점을 포괄하는 최초의 교차 장면, 다중 뷰 FloAff-Kitchen 벤치마크를 구축했습니다. 세 가지 벤치마크 환경에서 수행된 광범위한 실험 결과, 제안하는 방법이 강력한 기준 모델보다 일관되게 우수한 성능을 보이며, 부분 분석 연구를 통해 각 구성 요소의 기여도를 검증했습니다. 프로젝트 페이지: https://csu-hero-lab.github.io/FloAff-Kitchen_Web/
Mobile manipulation requires robots to identify Floor Affordance (FloAff) that maximizes downstream manipulation success rather than merely ensuring navigation feasibility. FloAff prediction is a target-conditioned local spatial reasoning problem, yet existing methods suffer from representation ambiguity caused by irrelevant spatial context and arbitrary object orientations, while entangling shared and task-specific knowledge across heterogeneous manipulation skills. To address these challenges, we propose a unified framework for FloAff prediction from egocentric multimodal perception, consisting of canonical representation learning and progressive affordance prior learning. Specifically, we introduce a Canonical Floor Affordance Representation (CFAR), which learns canonical interaction geometry by preserving affordance-relevant local structure while eliminating nuisance spatial variations unrelated to robot base placement. We further propose Progressive Floor Affordance Learning (PFAL), which learns transferable FloAff priors from a foundation manipulation task and progressively adapts them to heterogeneous downstream manipulation skills. To facilitate systematic evaluation, we establish the first cross-scene, multi-view FloAff-Kitchen benchmark covering diverse manipulation skills, scene layouts, furniture styles, and viewpoints. Extensive experiments on three benchmark settings demonstrate that our method consistently outperforms strong baselines, while ablation studies validate the contribution of each proposed component. Project page: https://csu-hero-lab.github.io/FloAff-Kitchen_Web/
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.