2607.27518v1 Jul 29, 2026 cs.AI

에이전트 기반 벤치마크의 결함 감지를 위한 자동화된 트랜스크립트 분석

Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks

Magda Dubois
Magda Dubois
Citations: 419
h-index: 7
Nelson Gardner-Challis
Nelson Gardner-Challis
Citations: 4
h-index: 1
H. Coppock
H. Coppock
Citations: 463
h-index: 11
Jeff T Mohl
Jeff T Mohl
Citations: 115
h-index: 3
Benjamin Allan-Rahill
Benjamin Allan-Rahill
Citations: 0
h-index: 0
Kaelan Yim
Kaelan Yim
Citations: 0
h-index: 0
Damian S'ojka
Damian S'ojka
Citations: 0
h-index: 0
James Mann
James Mann
Citations: 0
h-index: 0
Justin Olive
Justin Olive
Citations: 5
h-index: 1

최첨단 모델의 성능은 종종 에이전트 기반 벤치마크를 사용하여 평가됩니다. 이러한 결과를 신뢰하려면 벤치마크가 주장하는 내용을 정확하게 측정하고, 유효성을 저해하는 결함이 없어야 합니다. SWE-Bench-Verified와 같은 벤치마크에 대한 이전의 수동 감사에서 트랜스크립트 내 여러 가지 유효성 문제를 발견했습니다. 그러나 수동 검토는 확장하기 어렵고, 자동화된 방법이 벤치마크 유효성을 손상시키는 결함을 안정적으로 찾아낼 수 있을지 불확실합니다. 본 논문에서는, 진실 데이터 접근, 도구 오류, 추론 취약점 및 답변 형식의 모호성 등 네 가지 유형의 유효성 문제를 감지하기 위한 AI 스캐너를 개발했습니다. 각 유형에 대한 채점 기준을 작성하여 인간 레이블링을 안내하고, Inspect Evals 벤치마크의 보류된 테스트 세트에 대해 스캐너를 인간 레이블과 비교 평가했습니다. 당사의 스캐너는 광범위하게 사용되는 다섯 가지 벤치마크에서 여러 가지 유효성 문제를 확인했으며, 이는 무작위 수동 검사로는 발견하기 어려운 경우도 포함되었습니다. 모든 문제가 식별되지는 않았으며, 스캐너의 성능은 벤치마크, 기준 및 모델에 따라 크게 달랐습니다. 당사는 평가 분야의 더 광범위한 표준화 부족으로 인해 스캐너 성능이 저하되는 등, 더 강력한 품질 보증을 위한 해결해야 할 과제를 강조합니다. 전반적으로 이러한 결과는 자동화된 트랜스크립트 분석을 사용하여 벤치마크 품질을 보다 포괄적으로 감사하는 데 사용할 수 있음을 입증하는 사례 연구입니다.

Original Abstract

Capabilities of frontier models are often assessed using agentic benchmarks. To trust these results, benchmarks must accurately measure what they claim to and be free from invalidating flaws. Previous manual audits of benchmarks such as SWE-Bench-Verified have uncovered several validity issues in transcripts. However, manual review is difficult to scale, and it is unclear whether automated methods can reliably surface flaws that compromise benchmark validity. In this paper, we developed AI scanners to detect four types of validity issues: ground truth access, tool failure, guessing vulnerability, and answer format ambiguity. We produced grading rubrics for each to instruct human labeling, and evaluated the scanners against human labels on a held-out test set of Inspect Evals benchmarks. Our scanners identified several verified quality issues in five widely used benchmarks, including cases unlikely to be caught by random manual inspection. Not all cases were identified, and scanner performance varied substantially across benchmarks, criteria and models. We highlight several open challenges to be addressed to improve scanners for stronger quality assurance claims, including broader standardization gaps in the evaluation field that degrade scanner performance. Together, these results serve as a proof of concept for using automated transcript analysis to audit benchmark quality more broadly.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!