ATOM-Bench: 조작 정책의 기본 능력 및 조합적 일반화 성능을 평가하기 위한 실제 환경 벤치마크
ATOM-Bench: A Real-World Benchmark for Atomic Skills and Compositional Generalization in Manipulation Policies
일반적인 조작 정책은 로봇 제어의 기초 모델로 점점 더 많이 제시되고 있지만, 실제 환경에서의 일반화 성능을 진단하는 것은 여전히 어렵습니다. 특정 작업에서는 성공할 수 있지만, 세분화된 기본 동작을 수행하거나 학습된 기술을 새로운 작업 구조에 결합하는 데 실패할 수 있습니다. 본 논문에서는 조작 정책의 기본 능력과 조합적 일반화 성능을 평가하기 위한 실제 환경 벤치마크인 **ATOM-Bench**를 소개합니다. ATOM-Bench는 테이블 위 조작 작업을 모터 원자와 명령어 원자로 분해하고, 단일 팔 로봇 및 이중 팔 로봇 트랙에 걸쳐 총 30개의 기본 작업과 24개의 숨겨진 조합 작업으로 구성되어 있습니다. 우리는 3,000개의 인간 시연 데이터를 수집하여 기본 동작 미세 조정에 사용하고, 재현 가능한 실제 환경 평가를 지원하기 위해 시연 데이터와 평가 실행 데이터를 모두 공개합니다. 정책은 기본 작업에 대해 미세 조정되며, 기본 동작 습득 능력과 숨겨진 조합 작업 성능을 평가합니다. 또한, 약한 기본 능력으로 인한 실패와 제한적인 조합 재사용으로 인한 실패를 구별하기 위해 Atomic Score (AS) 및 Compositional Failure Share (CFS)를 도입했습니다. 5가지 대표적인 조작 정책에 대해 총 2,700번의 실제 실행 테스트를 수행한 결과, 현재의 정책은 간단한 명령어 기반 기술을 습득할 수 있지만, 여전히 세분화된 모터 동작, 계산 및 논리적 필터링에는 어려움을 겪습니다. 더욱 중요한 점은, 뛰어난 기본 능력 성능이 항상 숨겨진 조합 작업으로 잘 이어지는 것은 아닙니다. ATOM-Bench는 약한 모터 실행, 불량한 명령어 연결 또는 제한적인 조합 재사용으로 인해 발생하는 문제점을 진단할 수 있는 테스트 환경을 제공합니다.
Generalist manipulation policies are increasingly presented as foundation models for robotic control, but their real-world generalization remains difficult to diagnose. A policy may succeed on demonstrated tasks while still failing to execute fine-grained atomic skills or recombine learned skills in new task structures. We introduce \textbf{ATOM-Bench}, a real-world benchmark for evaluating both atomic skills and compositional generalization in manipulation policies. ATOM-Bench factorizes tabletop manipulation into motor atoms and instruction atoms, and contains 30 atomic tasks and 24 held-out compositional tasks across paired single-arm and dual-arm robot tracks. We collect 3,000 human demonstrations for atomic fine-tuning and release both the demonstration data and evaluation rollout data to support reproducible real-world evaluation. Policies are fine-tuned on atomic tasks and evaluated on both atomic skill acquisition and held-out compositional tasks. We further introduce Atomic Score (AS) and Compositional Failure Share (CFS) to distinguish failures caused by weak atomic skills from failures caused by limited compositional reuse. Through 2,700 physical rollouts on five representative manipulation policies, we find that current policies can acquire simple instruction-grounding skills, but still struggle with fine-grained motor atoms, counting, and logical filtering. More importantly, strong atomic performance does not reliably transfer to held-out compositional tasks. ATOM-Bench provides a diagnostic testbed for studying whether failures arise from weak motor execution, poor instruction grounding, or limited compositional reuse.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.