2606.31200v1 Jun 30, 2026 cs.AI

자율적인 RAG-VLM: 자기 성찰적 계획을 통한 수동성 인식 기반의 검색 증강 생성 모델로 로봇 집기 성능 향상

Agentic RAG-VLM: Affordance-Aware Retrieval-Augmented Generation with Self-Reflective Planning for Robotic Grasping

Jiaxu Wang
Jiaxu Wang
Citations: 285
h-index: 10
Tao Chen
Tao Chen
Citations: 307
h-index: 6
Lizheng Liu
Lizheng Liu
Citations: 0
h-index: 0
Ruiqi Tian
Ruiqi Tian
Citations: 0
h-index: 0
JiGuang Huo
JiGuang Huo
Citations: 0
h-index: 0
Zhongxue Gan
Zhongxue Gan
Citations: 160
h-index: 5
Ziyue Jiang
Ziyue Jiang
Citations: 0
h-index: 0

복잡한 환경에서 일반화된 로봇 집기는 인간이 사용하는 비정형 공간에 로봇 팔을 배치하는 데 필수적입니다. 그러나 기존의 VLM(Vision-Language Model, 시각 언어 모델) 기반 방법은 객체 매칭을 위해 시각적 유사성에 의존하며, 손잡이의 잡기 용이성이나 재료의 파괴되기 쉬운 정도와 같은 물리적인 특성을 고려하지 않습니다. 또한 공간 추론 또는 실패 복구 없이 개방형 루프로 작동하여, 물체가 밀집되거나 다양한 물리적 특성을 가진 경우 성능이 제한됩니다. 본 논문에서는 VLM 기반의 의미 이해와 물리적으로 기반된 집기 실행을 통합하는 통일된 프레임워크인 Agentic RAG-VLM을 제시합니다. Agentic RAG-VLM은 검색 증강 생성(RAG)과 시각 언어 모델(VLM), 그리고 자율적인 자기 성찰적 계획을 결합하여 작동하며, 다음 세 가지 핵심 구성 요소로 이루어져 있습니다: (1) 계층적 수동성 인식 기반 RAG(HAA-RAG): 이는 유형, 재료, 파괴되기 쉬운 정도 및 잡기 가능한 영역을 포함하는 4차원 수동성 설명자를 인코딩하고, 시각적인 특징 대신 기능적인 수동성 호환성을 기준으로 전략을 검색합니다; (2) 장면 그래프 제약 추론기: VLM의 인식 결과를 바탕으로 공간 관계 그래프를 생성하고, 근접성, 가려짐 및 지지 제약을 구체적인 집기 파라미터 조정으로 변환합니다; (3) 자율적인 자기 성찰적 파이프라인: 14가지 유형의 실패 분류 체계와 세 단계의 적응형 재시도 메커니즘을 통해 폐쇄 루프 기반의 집기 성능을 개선합니다. 단일 집기, 상호 작용 및 장기 시나리오를 포함하는 12개의 작업으로 구성된 벤치마크에서 총 360번의 실험을 수행한 결과, Agentic RAG-VLM은 78.3%의 성공률을 달성했으며, 이는 VLM만을 사용한 기준 모델보다 53.3%p 더 높은 성능입니다. 이는 수동성 인식 기반 검색, 장면 그래프 추론 및 자율적인 복구 메커니즘이 견고한 조작을 위해 공동으로 필수적임을 보여줍니다.

Original Abstract

Generalizable robotic grasping in cluttered environments is essential for deploying manipulators in unstructured human spaces, yet existing VLM-based methods rely on visual similarity for object matching, neglecting physical affordances such as handle graspability and material fragility, and operate open-loop without spatial reasoning or failure recovery, limiting their effectiveness when objects are densely packed or physically diverse. We present Agentic RAG-VLM, a unified framework that bridges VLM-based semantic understanding and physically grounded grasp execution by integrating retrieval-augmented generation (RAG) with vision-language models (VLMs) and agentic self-reflective planning. Agentic RAG-VLM introduces three tightly coupled components: (1) a Hierarchical Affordance-Aware RAG (HAA-RAG) that encodes four-dimensional affordance descriptors, including type, material, fragility, and graspable region, and retrieves strategies by functional affordance compatibility rather than visual appearance; (2) a Scene Graph Constraint Reasoner that constructs spatial relationship graphs from VLM perception and translates proximity, occlusion, and support constraints into concrete grasp parameter adjustments; and (3) an Agentic Self-Reflective Pipeline with a 14-type failure taxonomy and three-level adaptive retry for closed-loop grasp refinement. Evaluated on a 12-task benchmark spanning single-grasp, interactive, and long-horizon scenarios with 360 trials per configuration, Agentic RAG-VLM achieves 78.3 percent overall success, a 53.3 percentage-point absolute gain over VLM-only baselines, demonstrating that affordance-aware retrieval, scene graph reasoning, and agentic recovery are jointly essential for robust manipulation.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!