로보윗츠: 로봇의 창의적 문제 해결에 대한 예상치 못한 도전 과제
RoboWits: Unexpected Challenges for Robotic Creative Problem Solving
실제 환경에서 작동하는 로봇에게는 추론, 적응 능력 및 예기치 않은 상황에서의 창의적인 문제 해결 능력이 필수적입니다. 그러나 현재 로봇 벤치마크는 주로 숙련도 실행에 중점을 두고 있으며, 이러한 인지적 추론 능력에 대한 제한적인 정보를 제공합니다. 본 연구에서는 인지적 추론, 창의적인 도구 사용 및 예기치 못한 상황에 대한 견고성을 체계적으로 평가하기 위해 설계된 양팔 로봇 벤치마크인 '로보윗츠(RoboWits)'를 소개합니다. 고품질의 추론 중심적인 예기치 못한 시나리오를 확장 가능하게 구축하기 위해, 본 연구에서는 시드 작업 생성 및 검증, 지표 생성, 장면 생성 및 작업 변형을 위한 다중 에이전트 협업 프레임워크로 구성된 자동화된 작업 생성 파이프라인을 제안합니다. 이 파이프라인을 사용하여 30개의 다양한 시드 작업을 선별하고 기하학, 재료 및 조립 기반 추론에 따른 난이도가 등급화된 208개의 변형 작업을 생성했습니다. 인기 있는 로봇 정책, 사전 학습된 VLA(Visual Language Agent) 및 오라클 상태 계획기를 사용하여 성능을 평가했습니다. 그 결과, 상당한 성능 격차가 나타났습니다. 사전 학습된 VLA는 단일 작업에 대한 미세 조정 후 시드 작업에서는 어느 정도 성공을 거두었지만, 변형된 작업에서는 어려움을 겪었는데, 이는 추론, 전략 적응 및 기만적이거나 제약적인 환경에 대한 견고성이 필요한 조작 작업에서 이러한 모델의 취약성을 시사합니다. 프로젝트 페이지는 https://umass-embodied-agi.github.io/RoboWits 에서 확인할 수 있습니다.
The ability to reason, adapt, and creatively solve problems under unexpected challenges is essential for robots operating in real-world environments. However, current robotic benchmarks primarily emphasize skill-level execution and provide limited insight into such cognitive reasoning capabilities. We introduce RoboWits, a bi-manual robotic benchmark designed to systematically evaluate cognitive reasoning, creative tool use, and robustness to unexpected conditions. To enable scalable construction of high-quality reasoning-centric unexpected scenarios, we propose an automated task generation pipeline formulated as a multi-agent cooperative framework, comprising agents for seed task generation and verification, metric generation, scene generation, and task mutation. Using the pipeline, we curated 30 diverse seed tasks and 208 tasks with mutations and graded difficulty across geometry, material, and assembly-based reasoning. We benchmark popular robot policies, pre-trained VLAs, and oracle-state planners. Our results reveal a significant performance gap: while pre-trained VLAs exhibit preliminary success on seed tasks after single-task fine-tuning, they struggle to perform on mutated tasks, implying their brittleness in manipulation tasks requiring reasoning, strategy adaptation, and robustness to deceptive or constrained environments. Project page is available at https://umass-embodied-agi.github.io/RoboWits.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.