오프라인 선호도 기반 경로 평가
Offline Preference-Based Trajectory Evaluation
에이전트 시스템의 오프라인 평가는 종종 경로를 최종 성공 여부로 단순화하여 부분적인 진행 상황에 대한 정보를 버리고 광범위한 동점 결과를 초래하며, 이는 효과적인 샘플 크기를 줄이고 시스템을 구별하는 능력을 약화시켜 상당한 통계적 비효율성을 야기합니다. 본 연구에서는 경로의 진행 상황 및 회복 시간 프로필에 대한 시간적 선호도를 직접적으로 비교하는 선호도 기반 경로 평가 방법을 제안합니다. 다양한 에이전트 및 인터랙티브 벤치마크에서, 표준적인 성공 기반 지표는 대략 75%의 경우 동점 결과를 나타내는 반면, 경로 정보를 고려한 선호도는 동점을 약 35%로 줄여 구별력, 순위 안정성 및 데이터 효율성을 향상시킵니다. 이러한 결과는 벤치마크 포화 현상이 종종 수집된 데이터 품질이나 문제 난이도 때문이라고 여겨지지만, 평가 지표의 선택 또한 중요한 원인일 수 있음을 시사합니다.
Offline evaluation of agentic systems often collapses trajectories to terminal success, discarding information about partial progress and inducing widespread ties, creating substantial statistical inefficiency by reducing effective sample size and weakening the ability to distinguish systems. We propose preference-based trajectory evaluation, which compares trajectories directly through temporal preferences over progress and time-to-return profiles. We find that, across diverse agentic and interactive benchmarks, standard success-based metrics produce tied comparisons on roughly 75% of instances, whereas trajectory-aware preferences reduce ties to roughly 35%, improving discriminative power, ranking stability, and data efficiency. Our results suggest that benchmark saturation, often attributed to poor data collection or problem difficulty, may also be explained by the choice of evaluation measure.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.