APIVOT: 적응적 계획을 위한 상호 연결된 시각-언어 사고
APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts
장기적인 로봇 계획은 의미론적 작업 구조와 기하학적 실행 가능성에 대한 통합적인 추론을 요구합니다. 작업을 성공적으로 수행하기 위해서는 로봇이 목표를 분해하고, 작업과 관련된 객체를 선택하며, 행동 순서를 결정해야 합니다. 동시에 공간 제약 조건, 예를 들어 제한된 자유 공간 및 객체 충돌을 만족하는 계획을 수립해야 합니다. 본 연구에서는 장기적인 계획을 위해 언어와 시각적 사고를 적응적으로 결합하는 VLM 기반의 플래너인 APIVOT를 제안합니다. APIVOT는 언어를 활용하여 의미론적 추론을 수행하고, 동시에 기하학적 실행 가능성을 내부적으로 검증하기 위해 시각적 정보를 미래 상태로 사용합니다. 장기적인 주방 작업 환경에서 APIVOT는 범용 VLM 및 기존 계획 프레임워크보다 뛰어난 성능을 보이며, 특히 공간 제약이 강한 환경에서 가장 큰 개선 효과를 나타냅니다. 연구 결과, APIVOT는 의미 있는 모드 선택 행동을 학습하며, 이는 시각-언어적 사고의 적응적인 결합이 계획 성공률과 추론 효율성을 모두 향상시킨다는 것을 보여줍니다.
Long-horizon robot planning requires jointly reasoning over semantic task structure and geometric feasibility. To successfully execute a task, a robot must decompose goals, select task-relevant objects, and sequence actions, while ensuring that plans satisfy spatial constraints such as limited free space and object collisions. In this work, we propose APIVOT, a VLM-based planner that adaptively interleaves language and visual thoughts for long-horizon planning. APIVOT learns to leverage language for semantic reasoning, while using visual thoughts as imagined future states for internal verification of geometric feasibility. On long-horizon kitchen tasks, APIVOT outperforms general-purpose VLMs and prior planning frameworks, achieving the largest gains in spatially constrained settings. We find that APIVOT learns meaningful modality selection behavior, demonstrating that adaptive interleaving of vision-language thoughts improves both planning success and reasoning efficiency.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.