OpenClawBench: 실제 에이전트 실행 경로에서의 프로세스 측 이상 현상 벤치마킹
OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajectories
작업 성공은 실제 에이전트 실행 과정에서 발생하는 프로세스상의 이상을 가릴 수 있습니다. 에이전트는 최종 작업 검증 단계를 통과하더라도 여전히 해결되지 않은 모호성, 안전하지 않은 외부 쓰기, 무시된 오류, 취약한 기반의 약속, 또는 능력 범위를 벗어난 과도한 기능 사용 등의 문제를 야기할 수 있습니다. 우리는 이러한 불일치를 '결과-프로세스 격차(Outcome-Process Gap)'라고 정의하고, 실제 에이전트 실행 과정에서 발생하는 프로세스 측 이상 현상을 측정하고 관리하기 위한 대규모 데이터셋인 OpenClawBench를 소개합니다. OpenClawBench는 6개의 소스 모델에서 생성된 BFCL 기반의 OpenClaw 세션으로 구성되었으며, 31,264개의 주석이 달린 실행 경로 데이터를 포함합니다. 이 데이터셋은 작업 검증 결과와 구조화된 프로세스 증거를 연결합니다. FullTax는 이러한 연결된 실행 경로를 사용하여 구조화된 이상 현상 지도(anomaly supervision)를 생성하며, 여기에는 이진 레이블, 근거 증명, 시작/범위 위치 정보, 심각도, 복구 가능성 및 5가지 수준의 이상 현상 분류가 포함됩니다. OpenClawBench를 통해 우리는 결과-프로세스 격차를 측정할 수 있게 되었습니다. 31,135개의 작업 검증을 통과한 실행 경로 중에서, FullTax에 따르면 2,904개가 여전히 프로세스상의 이상 현상을 보이는 것으로 분류되었습니다. 이러한 결과는 성공만을 기준으로 평가하는 것이 실제 에이전트 실행 과정에서 발생하는 특정 유형의 프로세스 오류를 놓칠 수 있음을 보여줍니다. 고신뢰도 FullTax 데이터셋을 사용하여 학습된 LoRA-fine-tuned Gemma 3 12B 검출기는 더 깨끗한 레이블로 구성된 테스트 데이터 세트에서 이진 F1 점수가 0.729를 달성했습니다. OpenClawBench는 실제 에이전트 실행 로그를 감사 가능하고 재사용 가능한 지도 학습 데이터로 변환하여, 런타임 에이전트의 안정성을 연구, 진단 및 운영적으로 모니터링하는 데 활용될 수 있습니다.
Task success can hide process anomalies in real-world agent executions. An agent may pass the final task oracle while still accumulating unresolved ambiguity, unsafe external writes, ignored errors, weakly grounded commitments, or capability-boundary overcommitment. We study this mismatch as the Outcome-Process Gap and introduce OpenClawBench, a large-scale dataset for measuring and supervising process-side anomalies in real agent execution processes. OpenClawBench is built from BFCL-driven OpenClaw sessions produced by 6 source models and contains 31,264 annotated trajectories. It aligns task-oracle outcomes with structured process evidence. FullTax converts the aligned trajectories into structured anomaly supervision: binary labels, supporting evidence, onset/span localization, severity, recoverability, and a 5-class anomaly taxonomy. Using OpenClawBench, we make the Outcome-Process Gap measurable. Among 31,135 oracle-passing executions, 2,904 are still labeled process-anomalous under FullTax. These results show that success-only evaluation misses a concrete class of process-side failures in real agent executions. A LoRA-fine-tuned Gemma 3 12B detector trained on the high-confidence FullTax supervised pool reaches binary F1=0.729 on the cleaner-labels held-out test split. Together, OpenClawBench turns real agent execution logs into auditable and reusable supervision for studying, diagnosing, and operationally monitoring runtime agent reliability.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.