Embodied-BenchClaw: 구체화된 공간 지능 벤치마크 구축을 위한 자율적 다중 에이전트 시스템
Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction
벤치마크는 구체화된 공간 지능을 평가하는 데 필수적이지만, 그 구축은 노동 집약적이고 재사용하기 어렵고 유지 관리가 어렵습니다. 기존의 구체화된 벤치마크는 종종 정적이어서 모델이 발전함에 따라 빠르게 포화 상태가 되어 새로운 기능을 구분하는 능력을 제한할 수 있습니다. 본 연구에서는 구체화된 공간 지능 벤치마크를 구축하기 위한 자율적인 에이전트 시스템인 Embodied-BenchClaw를 제안합니다. 사용자가 지정한 평가 목표를 기반으로, Embodied-BenchClaw는 다섯 단계의 파이프라인을 통해 완전하고 지속적으로 업데이트 가능한 벤치마크 패키지를 자동으로 생성합니다. 이 파이프라인은 계획, 구축 및 평가를 담당하는 세 개의 에이전트가 조정하여 운영됩니다. 재사용성과 신뢰성을 향상시키기 위해 Embodied-BenchClaw는 확장 가능한 스킬 라이브러리와 프로세스 품질 관리를 도입하여 벤치마크 구축을 구성 가능하고 검증 가능하며 수정 가능하게 만듭니다. 우리는 실내 공간 추론, 야외 공간 추론, 로봇 조작, 사족 보행 로봇 내비게이션, UAV/항공뷰 이해 및 기존 벤치마크 개선 등을 포괄하는 여러 벤치마크를 구현했습니다. 이러한 벤치마크는 다양한 구체화된 플랫폼, 데이터 소스 및 공간 능력을 포함합니다. 인간 평가, 심사위원 기반 평가, 일관성 검사, 비용 분석 및 부분 제거 실험을 통해 Embodied-BenchClaw가 수동 노력을 줄이고 검증 가능하고 실행 가능하며 유지 관리 가능하고 진단적으로 유용한 구체화된 공간 벤치마크를 구축할 수 있음을 보여줍니다.
Benchmarks are essential for evaluating embodied spatial intelligence, yet their construction is labor-intensive, hard to reuse, and difficult to maintain. Existing embodied benchmarks are often static and may quickly become saturated as models improve, limiting their ability to distinguish new capabilities. We propose Embodied-BenchClaw, an autonomous agentic system for constructing embodied spatial intelligence benchmarks. Given a user-specified evaluation intent, Embodied-BenchClaw automatically produces a complete and continually updatable benchmark package through a five-stage pipeline: intent blueprinting, data collection, structuring and cleaning, benchmark synthesis, and evaluation reporting. The pipeline is coordinated by three agents for planning, construction, and evaluation. To improve reusability and reliability, Embodied-BenchClaw introduces an extensible Skill Library and process quality control, enabling benchmark construction to be composable, verifiable, and repairable. We instantiate multiple benchmarks covering indoor spatial reasoning, outdoor spatial reasoning, robotic manipulation, quadruped robot navigation, UAV/aerial-view understanding, and static benchmark enhancement. These benchmarks span diverse embodied carriers, data sources, and spatial capabilities. Experiments with human evaluation, judge-based assessment, consistency checks, cost analysis, and ablations show that Embodied-BenchClaw can construct verifiable, executable, maintainable, and diagnostically useful embodied spatial benchmarks with reduced manual effort.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.