ClawTrack: 실제 환경 자율 에이전트의 추적 수준 평가 및 개선을 위한 연구
ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents
LLM 기반 에이전트가 복잡하고 다단계 워크플로우에 적용됨에 따라, 중요한 평가 격차가 발생했습니다. 대부분의 기존 벤치마크는 최종 결과만을 평가하며, 신뢰할 수 있는 추론과 운 좋은 성공을 구별하거나 특정 프로세스 결함을 지적하기 어렵기 때문에 장기적인 작업에서 원인 분석이 어렵습니다. 본 연구에서는 ClawTrack이라는 이중 평가 벤치마크를 제시합니다. ClawTrack은 에이전트가 무엇을 달성했는지 (Task Score)와 어떻게 달성했는지를 동시에 측정합니다. ClawTrack은 8개의 도메인에 걸쳐 320개의 작업과 25개 이상의 결정적 모의 서비스를 포함합니다. Process Grader는 각 추론 단계를 4가지 차원(목표 일치성, 효율성, 정보 활용, 결과 검증)을 기준으로 평가하며, 각 작업별로 12,541개의 세부 기준 항목을 사용합니다. 21개의 모델을 16,000회 이상의 테스트를 통해 분석한 결과, 다음과 같은 사실을 확인했습니다: (1) 프로세스 점수는 성공과 실패를 특정 추론 차원에 연결하여 평가함으로써, 결과 중심의 평가로는 파악할 수 없는 운 좋은 결과를 걸러낼 수 있습니다; (2) 4가지 차원은 상호 보완적이며, 특히 결과 검증이 체계적인 병목 현상입니다; (3) 이 프레임워크는 다양한 LLM 평가기 사이에서도 일관성을 유지합니다; (4) 프로세스 기반의 경로 필터링은 모델 크기에 관계없이 훈련 후 성능 개선에 일관된 효과를 가져옵니다.
As LLM-based agents are deployed in complex, multi-step workflows, a critical evaluation gap has emerged: most existing benchmarks judge only final outcomes, unable to distinguish reliable reasoning from lucky success or attribute failures to specific process deficiencies, hindering attribution in long-horizon tasks. In this work, we present ClawTrack, a dual-assessment benchmark that simultaneously measures what an agent achieves (Task Score) and how it achieves it (Process Score). ClawTrack comprises 320 tasks across 8 domains with 25+ deterministic mock services. A Process Grader scores each reasoning turn along four dimensions (goal alignment, efficiency, information utilization, and result verification), anchored by 12,541 task-specific rubric items. Evaluating 21 models over 16,000+ trials, we find that: (1) process scores effectively attribute success and failure to specific reasoning dimensions, filtering lucky passes invisible to outcome-only evaluation; (2) the four dimensions are complementary, with result verification as the systematic bottleneck; (3) the framework is robust to evaluator choice across different judge LLMs; and (4) process-based trajectory filtering yields consistent post-training improvements across model scales.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.