2606.17005v1 Jun 15, 2026 cs.AI

최첨단 AI 평가의 공개 아카이브를 위한 베이지안 추론 및 의사 결정 감사

Bayesian Inference and Decision Audits for Public Archives of Frontier AI Evaluations

Yanan Long
Yanan Long
Citations: 104
h-index: 4

공개 AI 평가는 종종 최종 순위표로 해석되지만, 그 기반에는 보고 규칙, 벤치마크 수정 및 누락된 데이터에 의해 형성된 선택적인 시계열 데이터가 존재합니다. LiveBench 및 Open LLM Leaderboard v2의 반복적인 공개 아카이브는 주요 장기 기록을 제공하며, LMArena는 선호도 스트레스 테스트를 수행하고, GAIA와 tau-bench는 제한적인 에이전트 기반 실험에 기여합니다. 이러한 아카이브들은 베이지안 추론 문제를 야기하는데, 특정 보고 규칙 하에서 1,000개 이상의 시스템으로 구성된 최종 결과 예시는 두 가지 이전 기록과 일치하며, 동일한 최종 모델 하에서 최대값에서 0.05 이내의 값을 얻는 데 걸리는 시간이 23.03 또는 75.13이 됩니다. 합성 후처리 비교 분석에서는 관찰 환경에 따라 행동 관련 진단 결과가 다르게 나타납니다. 후보 선택을 고려한 최첨단 모델은 합성 데이터 복구, 객체 아카이브 예측, 선호도 전송 및 불확실성 보정에 실패하며, 이에 따라 고정된 감사 기준은 해당 모델의 과장된 주장을 거부합니다. 아카이브 및 심사 프로토콜은 공개 평가 기록을 재구성하고, 검증된 시간 경계를 설정하며, 근거 없는 최첨단 주장을 거짓으로 판별합니다.

Original Abstract

Public AI evaluations are often read as terminal leaderboards, yet the underlying evidence is a selective time series shaped by reporting rules, benchmark revisions, and missingness. Repeated public archives for LiveBench and Open LLM Leaderboard v2 serve as the primary longitudinal record; LMArena provides a preference stress test; and GAIA and tau-bench contribute limited agentic pilots. Together, these archives instantiate a Bayesian inference problem: under a fixed reporting convention, one constructed terminal-only example over $1{,}000$ systems is compatible with two pre-terminal histories, yielding times of $23.03$ or $75.13$ to reach within $0.05$ of the ceiling under the same terminal-tail model. In synthetic posterior comparisons, action-facing diagnostics differ across observation regimes. The candidate selection-aware frontier model fails synthetic recovery, objective-archive prediction, preference transfer, and uncertainty calibration; correspondingly, fixed audit gates reject its stronger claims. An archive-and-adjudication protocol reconstructs public evaluation histories, isolates a verified timing boundary, and falsifies unsupported frontier claims.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!