SciCode-Verified: 벤치마크 결함이 언어 모델의 과학 코딩 능력 저평가에 미치는 영향
SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models
SciCode는 언어 모델의 과학 코딩 능력을 평가하는 표준 지표로, 최첨단 과학 이론과 이를 실제 수치 코드로 구현하는 연구 수준의 문제를 다룹니다. 이는 인공지능 분석 지수(Artificial Analysis Intelligence Index)의 구성 요소이며, 정부 및 국가 연구소에서 정기적으로 활용되는 평가 도구입니다. 그러나 최근 SciCode 점수가 정체되는 현상이 나타났습니다. 가장 뛰어난 2026년 모델들은 약 60% 수준의 부분 문제 정확도를 보이며, 후속 모델 또한 전작과 비슷한 성능을 보였습니다. 본 연구는 이러한 정체의 원인이 SciCode 벤치마크 자체의 결함에 있음을 밝히고, 이를 해결하고자 합니다. 모든 65개 테스트 문제에 대한 전문가 검토 결과, 263개의 결함이 발견되었으며, 이 중 192개가 전체 주요 문제의 91%에 걸쳐 나타나, 정확하고 지시를 잘 따르는 솔루션조차도 잘못된 답변으로 채점되는 문제를 야기합니다. 이러한 문제는 재현 불가능한 정답, 지나치게 엄격한 허용 오차 또는 자체 모순적인 사양으로 인해 발생합니다. 특히 중요한 점은 이러한 성능 저하를 유발하는 결함의 78%가 단순한 문법 오류 검토가 아닌, 전문적인 물리학 또는 수학 지식을 통해서만 발견 가능하다는 것입니다. 우리는 확인 가능한 모든 결함을 수정하여 SciCode-Verified 버전을 만들었습니다. 이 과정에서 문제 정의 명확성 개선, 채점 방식 수정, 그리고 지나치게 관대한 테스트를 강화했습니다. 각 변경 사항은 그 이유와 함께 기록되었으며, 다른 전문가에 의해 독립적으로 재검토되었습니다. 수정된 벤치마크를 사용하여 최첨단 모델 12개의 성능을 재평가한 결과, 부분 문제 정확도가 45-60%에서 84-98%로, 주요 문제 정확도는 9-27%에서 69-92%로 크게 향상되었습니다. 이는 현재 최고 수준의 모델들이 SciCode가 제시하는 것보다 훨씬 뛰어난 과학 코딩 능력을 보유하고 있으며, 성능 저하의 원인은 모델 자체의 능력 부족이 아닌 평가 도구의 품질 문제에 있음을 시사합니다. 본 연구에서는 수정된 SciCode-Verified 버전을 전체 감사 기록과 함께 공개하여, 향후 표준으로 활용될 수 있도록 합니다.
SciCode is the standard measure of the scientific-coding ability of language models: research-level problems that demand both frontier scientific theory and its implementation as working numerical code. It is a component of the Artificial Analysis Intelligence Index and a standing evaluation in government and national-laboratory suites. Yet its scores have recently plateaued: the strongest 2026 models cluster tightly around 60\% subproblem accuracy, and a successor model ties its predecessor. We trace this stagnation to defects in the benchmark itself. A per-problem, domain-expert audit of all 65 test problems uncovers 263 defects; 192 of them, spread across 91\% of the main problems, cause correct, instruction-following solutions to be wrongly rejected---through non-reproducible gold answers, over-tight tolerances, or self-contradictory specifications. Critically, 78\% of these score-suppressing defects require specialized physics or mathematics knowledge to detect, not mere clerical proofreading. We corrected every confirmable defect to produce SciCode-Verified. The corrections add only the specifications a well-posed problem requires, repair grading, and tighten the tests that were too lenient; every change is recorded with its justification and independently re-checked by a second domain expert. We re-evaluate twelve frontier model snapshots on the corrected benchmark and find a substantial recovery: subproblem accuracy rises from 45--60\% to 84--98\%, and main-problem accuracy from 9--27\% to 69--92\%. State-of-the-art models are far more proficient in scientific coding than SciCode has suggested---the bottleneck was not model capability, but the quality of the evaluation instrument. We release SciCode-Verified with its complete audit trail as the corrected public standard.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.