2608.03219v1 Aug 04, 2026 cs.AI

도달 가능성이 실현 가능성을 의미하지 않는다: LLM 벤치마크 성능 향상의 근원을 추적하다

Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains

Jiaqing Xie
Jiaqing Xie
Citations: 54
h-index: 4
Ben Gao
Ben Gao
Citations: 231
h-index: 7
Tianfan Fu
Tianfan Fu
Citations: 7
h-index: 1
Yuqiang Li
Yuqiang Li
Citations: 29
h-index: 4
Yanbo Wang
Yanbo Wang
Citations: 453
h-index: 7
Wanhao Liu
Wanhao Liu
Citations: 119
h-index: 5
Yanchao Li
Yanchao Li
Citations: 16
h-index: 3

벤치마크 성능 향상은 종종 더 뛰어난 LLM의 능력을 나타내는 증거로 간주됩니다. 그러나 동일한 성능 향상이 모델 동작의 서로 다른 변화를 반영할 수 있습니다. 모델이 새로운 답변을 생성하거나, 이미 가능했던 답변을 생성할 수 있습니다. 집계 점수는 이러한 변화를 질문별로 구별하지 못합니다. 우리는 고정된 예산, 온도 및 답변 형식 하에서 질문 수준의 분석을 수행했습니다. '실현'이란 기본 배포 절차에서 올바른 답변이 생성되는 경우를 의미하며, '도달 가능'이란 특정 탐색 방법으로 지정된 예산 내에서 해당 답변을 찾을 수 있는 경우를 의미합니다. 먼저 추론 시간 레이어 라우팅이 도달 가능성을 확장할 수 있는지 테스트했습니다. 동일한 예산을 사용했을 때, 무작위 경로가 43개의 모델 및 작업 설정에서 구조화된 검색과 동등하거나 더 나은 결과를 보였습니다. 답변을 고려하지 않는 방법은 이러한 성능 향상의 대부분을 유지하지 못했으며, 이는 올바른 답변에 대한 접근이 필요합니다. 다음으로 도달 가능한 답변이 때로는 나타나지 않는 이유를 분석했습니다. 0.5B에서 31B까지의 규모의 6가지 경우를 분석한 결과, 특정 MLP 블록을 비활성화하면 미리 정의된 실패 사례의 68%에서 92%가 해결되었습니다. 다음으로 학습이 도달 가능성을 확장하여 격차를 해소하는지 테스트했습니다. 6번의 평가 중 5번에서 배포 성능은 향상되는 반면, 도달 가능한 최고 성능은 그대로 유지되거나 감소했습니다. DAPO의 경우, 배포 점수는 14.7점이 상승한 반면, 도달 가능한 최고 성능은 13.3점 하락했습니다. 따라서 분석한 모든 설정에서 실현 가능성과 도달 가능성은 항상 함께 변하지 않습니다. 능력 향상에 대한 주장은 동일한 평가 조건하에서 측정된 실현 성능과 도달 가능성을 모두 보고해야 합니다. 코드는 다음 링크에서 확인할 수 있습니다: https://github.com/LiZaiyuan0619/reachability-not-realization

Original Abstract

Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can reflect different changes in model behavior. A model may reach new answers, or produce answers that were already within reach. Aggregate scores do not distinguish these changes question by question. We establish a question-level audit under fixed budgets, temperatures, and answer formats. A question is realized when the default deployment procedure produces the correct answer. A question is reachable when a specified probe finds that answer within a fixed budget. We first test whether inference-time layer routing can expand reachability. Under a matched budget, random routes match or exceed structured search in all 43 model and task settings. Answer-blind procedures retain almost none of this gain, which instead requires access to the correct answer. We then ask why reachable answers sometimes fail to appear. Across six cases spanning 0.5B to 31B, silencing one identified MLP block repairs 68 to 92 percent of a predefined failure set. We next test whether training closes the gap by expanding reachability. In five of six matched evaluations, deployed performance rises while the reachable ceiling remains flat or falls. For DAPO, the deployed score rises by 14.7 points while the reachable ceiling falls by 13.3 points. Across the settings we audit, realization and reachability therefore do not always change together. Claims of capability expansion should report both realized performance and reachability under matched evaluation conditions. Code is available at https://github.com/LiZaiyuan0619/reachability-not-realization

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!