2608.02985v1 Aug 04, 2026 cs.LG

LLM 역테스팅에서의 시간적 누수: 측정, 검증 및 조정된 점수

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores

Bradly C. Stadie
Bradly C. Stadie
Citations: 2,378
h-index: 13
Zeyu Zhang
Zeyu Zhang
Citations: 19
h-index: 2

LLM 역테스팅에서 오염 여부를 확인하는 일반적인 방법은 학습 데이터 cutoff 시점 전후의 점수를 비교하는 것입니다. 본 연구에서는 이러한 방식이 유용한 정보를 제공하지 못함을 보여줍니다. 네 가지 주요 모델이 메모리화할 수 없는 질문에 대해 이 검증 방식을 통과하지 못했으며, 이는 cutoff 시점 이후에 해결된 모든 평가된 질문에서 나타났습니다. 그 이유는 구조적인 문제입니다. 모델은 원칙적으로 학습 cutoff 시점에 가까운 정보에 대해 더 많은 지식을 가지고 있으므로, 최신성이 누수와 유사하게 보이며, 수동적인 역테스팅으로는 이러한 두 가지 현상을 실제 능력과 구별할 수 없음을 증명합니다. 측정은 단순히 감지하는 것 이상으로, 역테스팅 외부의 정보를 필요로 합니다. 본 연구에서는 이를 두 가지 형태로 제공합니다. 알려진 cutoff 시점을 사용하면 경계에서의 누수를 식별할 수 있으며, 일치된 정상 컨트롤 그룹을 사용하면 전체적인 누수를 식별하고 누수 조정 점수를 얻을 수 있습니다. 또한 누수가 발생하는 위치를 분석한 결과, 누수는 집단적 지식과 상반되는 결과를 나타내고 학습 데이터에서 잘 다루어진 경우에 집중적으로 발생하며, 부분적인 암기 능력은 과도하게 보상받는 경향이 있음을 확인했습니다. 연구팀은 쌍둥이 모델에 인위적인 누수를 삽입하여 제안하는 추정 방법의 정확성을 검증했으며, 실제로 삽입된 누수의 양을 정확하게 복구하고 정상적인 질문에서는 0을 반환하는 것을 확인했습니다. 최첨단 모델에 적용한 결과, 특정 cutoff 시점에 국한된 특징을 감지했으며, 감사 성능의 최소 수준에서 단순히 최신 정보에 기반한 장점을 가진 것으로 보이는 다섯 개의 모델을 올바르게 식별했습니다. 역테스팅은 폐기할 필요가 없습니다. 단 하나의 신뢰할 수 있는 기준이 필요합니다.

Original Abstract

The standard check for contamination in LLM backtests is simple: compare scores before and after the training cutoff. We show this check is uninformative. Four flagship models fail it on questions they cannot have memorized: every scored question resolved after their cutoffs. The reason is structural. Models legitimately know more about times near their cutoff, so recency mimics leakage, and we prove no passive backtest can separate the two from genuine skill. Measurement, not just detection, requires information from outside the backtest. We supply it in two forms. A known cutoff identifies leakage at the boundary; a matched clean control identifies it globally and yields a leakage-adjusted score. We also derive where leakage hides: it concentrates on outcomes that surprised the crowd and were well covered in training, and partial memorization is disproportionately rewarded. We validate the estimators against ground truth by planting leakage in twin models, where they recover the injected dose and return null on clean questions. Deployed on frontier models, they detect one cutoff-localized signature and, at the audit's power floor, clear five models whose apparent advantages were recency alone. Backtests need not be discarded; they need one defensible reference.

0 Citations
0 Influential
6.5 Altmetric
32.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!