STAGE-Claw: 현실적인 시나리오를 위한 자동화된 상태 기반 에이전트 성능 평가
STAGE-Claw: Automated State-based Agent Benchmarking for Realistic Scenarios
대규모 언어 모델은 일상생활 애플리케이션을 위한 개인 비서(에이전트)의 핵심 기술로 점점 더 많이 사용되고 있지만, 이러한 에이전트를 평가하는 것은 여전히 어려운 과제입니다. 기존 벤치마크는 여전히 격리된 환경, 정적인 작업 설계, 그리고 단순한 점수 체계를 기반으로 하며, 이는 확장성을 저해하고 신뢰성 있는 개인 비서 평가를 향한 발전을 제한합니다. 본 논문에서는 STAGE-Claw라는 자동화 프레임워크를 소개합니다. STAGE-Claw는 상태 기반의 개인 컴퓨팅 환경에서 현실적인 개인 비서 시나리오를 구축하고 평가하는 데 사용됩니다. STAGE-Claw는 작업 힌트를 입력받아, 실제 환경, 작업 지시문, 정답 데이터, 그리고 관련 검증 프로그램을 자동으로 생성하고 검증합니다. 에이전트는 실제 운영 환경에서 평가되며, 최종 시스템 상태의 정확성을 기준으로 성능을 측정합니다 (텍스트 응답만으로는 평가하지 않습니다). 본 논문에서는 STAGE-Claw를 사용하여 40개의 도전적인 실생활 시나리오 기반 에이전트 작업을 포함하는 벤치마크를 구축하고, 11개의 최첨단 모델을 평가하며, 그들의 작업 점수, 비용, 도구 사용의 신뢰성, 그리고 일반적인 실패 패턴을 분석합니다. 전반적으로, STAGE-Claw는 실제 사용자 시나리오에서 에이전트를 평가하는 데 있어 확장 가능하고 상태 기반 접근 방식을 제공합니다.
Large language models are increasingly used to power personal agents for everyday applications, but evaluating these agents remains a challenge. Existing benchmarks still rely on sandboxed artifacts, static task design, and coarse scoring, which hinder scalability and limit progress toward reliable personal-agent evaluation. This paper introduces STAGE-Claw, an automated framework for building and evaluating realistic personal-agent scenarios in state-based personal-computing environments. Given a task hint, STAGE-Claw automatically creates and validates a realistic benchmark task with its environment, task prompts, ground truth, and related verification programs. Agents are then evaluated in realistic operating environments, where performance is measured by the correctness of the final system state rather than only the textual response. Using STAGE-Claw, this paper creates a benchmark with 40 challenging real scenario agent tasks, evaluates 11 frontier models, and analyzes their task scores, costs, tool-call reliability, and common failure patterns. Overall, STAGE-Claw offers a scalable, state-based way to evaluate agents in realistic user scenarios.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.