2607.28037v1 Jul 30, 2026 cs.LG

ClawTrack: 실제 환경 자율 에이전트의 추적 수준 평가 및 개선을 위한 연구

ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents

Linsen Guo
Linsen Guo
Citations: 175
h-index: 6
Xingchen Liu
Xingchen Liu
Citations: 76
h-index: 4
Jianing Wang
Jianing Wang
Citations: 19
h-index: 2
Xingjian Wu
Xingjian Wu
Citations: 969
h-index: 12
Xuhan Zhu
Xuhan Zhu
Citations: 23
h-index: 3
Xiaoyu Li
Xiaoyu Li
Citations: 113
h-index: 6
Xuezhi Cao
Xuezhi Cao
Citations: 47
h-index: 3
Xunliang Cai
Xunliang Cai
Citations: 231
h-index: 8
Junlin Liu
Junlin Liu
Citations: 393
h-index: 6

LLM 기반 에이전트가 복잡하고 다단계 워크플로우에 적용됨에 따라, 중요한 평가 격차가 발생했습니다. 대부분의 기존 벤치마크는 최종 결과만을 평가하며, 신뢰할 수 있는 추론과 운 좋은 성공을 구별하거나 특정 프로세스 결함을 지적하기 어렵기 때문에 장기적인 작업에서 원인 분석이 어렵습니다. 본 연구에서는 ClawTrack이라는 이중 평가 벤치마크를 제시합니다. ClawTrack은 에이전트가 무엇을 달성했는지 (Task Score)와 어떻게 달성했는지를 동시에 측정합니다. ClawTrack은 8개의 도메인에 걸쳐 320개의 작업과 25개 이상의 결정적 모의 서비스를 포함합니다. Process Grader는 각 추론 단계를 4가지 차원(목표 일치성, 효율성, 정보 활용, 결과 검증)을 기준으로 평가하며, 각 작업별로 12,541개의 세부 기준 항목을 사용합니다. 21개의 모델을 16,000회 이상의 테스트를 통해 분석한 결과, 다음과 같은 사실을 확인했습니다: (1) 프로세스 점수는 성공과 실패를 특정 추론 차원에 연결하여 평가함으로써, 결과 중심의 평가로는 파악할 수 없는 운 좋은 결과를 걸러낼 수 있습니다; (2) 4가지 차원은 상호 보완적이며, 특히 결과 검증이 체계적인 병목 현상입니다; (3) 이 프레임워크는 다양한 LLM 평가기 사이에서도 일관성을 유지합니다; (4) 프로세스 기반의 경로 필터링은 모델 크기에 관계없이 훈련 후 성능 개선에 일관된 효과를 가져옵니다.

Original Abstract

As LLM-based agents are deployed in complex, multi-step workflows, a critical evaluation gap has emerged: most existing benchmarks judge only final outcomes, unable to distinguish reliable reasoning from lucky success or attribute failures to specific process deficiencies, hindering attribution in long-horizon tasks. In this work, we present ClawTrack, a dual-assessment benchmark that simultaneously measures what an agent achieves (Task Score) and how it achieves it (Process Score). ClawTrack comprises 320 tasks across 8 domains with 25+ deterministic mock services. A Process Grader scores each reasoning turn along four dimensions (goal alignment, efficiency, information utilization, and result verification), anchored by 12,541 task-specific rubric items. Evaluating 21 models over 16,000+ trials, we find that: (1) process scores effectively attribute success and failure to specific reasoning dimensions, filtering lucky passes invisible to outcome-only evaluation; (2) the four dimensions are complementary, with result verification as the systematic bottleneck; (3) the framework is robust to evaluator choice across different judge LLMs; and (4) process-based trajectory filtering yields consistent post-training improvements across model scales.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!