끊임없이 변화하는 목표: 오픈 소스 채팅 LLM 출시 라인 전반에 걸친 신뢰도 측정 기준 점수 변동성에 대한 종단적 감사
The Moving Target: A Longitudinal Audit of Trust-Benchmark Score Drift Across Open-Source Chat LLM Release Lines
채팅-LLM 출시 라인에서 보고되는 신뢰도 측정 기준 점수는 동일한 라인의 여러 버전에서 그대로 유지되는 경우가 많으며, 이는 기본 모델이 릴리스 사이에 변경되지 않았다는 것을 암시합니다. 본 연구에서는 이러한 가정을 검증합니다. 우리는 Yi, Qwen, Mistral, 그리고 Gemma의 네 가지 오픈 소스 출시 라인을 세 번의 연속적인 공개 버전에 대해 감사했습니다. 각 버전은 200개의 항목으로 구성된 다섯 가지 채팅 평가 기준(TruthfulQA, BBQ, ToxiGen, CrowS-Pairs, 및 XSTest)을 사용하여 세 가지 프롬프트 템플릿 하에서 점수화되었습니다. 이 중 다섯 가지 신뢰도 측정 기준 변형 중 네 가지는 표준적이지 않으며, 그중 두 가지는 합성된 대리 지표입니다. 평균 절대 인접 버전 점수 변동률은 독립성을 기반으로 한 기준 null 값의 평균보다 몇 배 더 높습니다. 특정 신뢰도 측정 기준을 삭제하거나, 출시 라인을 제거하거나, 엄격한 점수를 적용하거나, 파라미터 크기가 일정하게 유지되는 경우에만 이 변동률이 감소합니다. 이러한 감사 환경에서 인용된 신뢰도 점수는 해당 버전의 점수에 국한됩니다. 새로운 릴리스가 나올 때마다 재측정되어야 하며, 이전 값을 그대로 사용할 수 없습니다. 폐쇄형 API, 더 큰 모델, 표준 프로토콜을 사용한 점수, 그리고 신뢰도 측정 기준 항목 부분 집합에 대한 불확실성은 본 연구의 범위를 벗어납니다.
Trust-benchmark scores reported on a chat-LLM release line are often carried across several checkpoints of the same line, as if the underlying model had not shifted between releases. We test that assumption. We audit four open-source release lines (Yi, Qwen, Mistral, and Gemma) at three successive public generations each. Each checkpoint is scored on a fixed 200-item basket of five chat-evaluation benchmarks: TruthfulQA, BBQ, ToxiGen, CrowS-Pairs, and XSTest, under three prompt templates. Four of the five benchmark variants are non-canonical, and two of those are synthetic proxies. The mean absolute adjacent-generation Score Drift Rate is several times the mean of an independence-based count-level reference null. It stays in the same band when we drop a benchmark, drop a release line, switch to strict scoring, or restrict to constant-parameter-size transitions. Within this audited setup, a quoted trust score should be treated as checkpoint-bound. It should be re-measured on each materially new release rather than carried forward. Closed APIs, larger models, canonical-protocol scores, and benchmark-item-subset uncertainty are out of scope.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.