2607.19104v1 Jul 21, 2026 cs.SE

SciCodePile: 도전적인 과학 코드 생성에 활용할 수 있는 128GB 규모의 데이터셋 및 실행 가능한 평가 도구

SciCodePile: A 128GB Corpus and Executable Benchmark for Challenging Scientific Code Generation

Jonathan Pan
Jonathan Pan
Citations: 67
h-index: 3
Jieke Shi
Jieke Shi
Citations: 1,005
h-index: 17
Weifeng Sun
Weifeng Sun
Citations: 125
h-index: 7
Ye Fan
Ye Fan
Citations: 39
h-index: 2
Gou Tan
Gou Tan
Citations: 29
h-index: 2
Yuan Yidi
Yuan Yidi
Citations: 6
h-index: 1
Swee Liang Wong
Swee Liang Wong
Citations: 57
h-index: 3
David Lo
David Lo
Citations: 104
h-index: 4
Yuchen Chen
Yuchen Chen
Citations: 410
h-index: 9

대규모 언어 모델(LLM)은 범용 코드 생성에서 뛰어난 성능을 보이지만, 과학 코드를 얼마나 잘 처리하는지는 아직 명확히 밝혀지지 않았습니다. 기존의 데이터셋과 벤치마크는 규모, 분야 범위 또는 실행 가능성 검증 측면에서 한계가 있어, 현재 LLM과 신뢰할 수 있는 과학 코드 생성기 간의 실제 격차를 제대로 평가하기 어렵습니다. 이러한 제한 사항을 해결하기 위해, 우리는 지금까지 가장 큰 규모의 과학 코드 데이터셋인 SciCodePile을 구축했습니다. 이는 37,737개의 공개 저장소에서 수집되었으며, 총 128GB의 코드로 구성되어 있으며, 다양한 계산 과학 분야를 포괄합니다. 이 데이터셋으로부터, 우리는 샌드박스 실행 환경과 자동 테스트 기능을 갖춘 200개의 작업으로 구성된 실행 가능한 벤치마크를 추가로 구축했습니다. 오픈 소스와 비공개 모델 모두에서 개발된 15개의 LLM을 사용하여 세 가지 작업(접두사-인접 문자열 완성, 중간 삽입, 실행 가능한 코드 생성)에 대한 성능을 평가했습니다. 결과는 과학 코드 생성이 여전히 매우 어려운 과제임을 보여줍니다. 가장 높은 CodeBLEU 점수는 두 가지 완성 작업에서 각각 38.13과 38.37에 불과했으며, 가장 강력한 모델조차 실행 가능한 벤치마크에서 Pass@1 기준으로 12.30%의 성능을 보였습니다. 이는 현재 모델이 신뢰할 수 있는 과학 코드 생성에 얼마나 멀리 떨어져 있는지 보여줍니다. SciCodePile의 학습 활용 가능성을 입증하기 위해, 데이터셋으로 추가 사전 훈련을 수행했을 때 과학 코드 완성 작업에서 CodeBLEU가 2.84배 향상되었으며, 데이터로 지시사항 기반 미세 조정을 수행했을 때 실행 가능한 벤치마크에서 Pass@1이 4.79배 향상된 것을 확인했습니다. 모든 코드와 데이터는 https://huggingface.co/SciCodePile 에서 확인할 수 있습니다.

Original Abstract

Large language models (LLMs) excel at general-purpose code generation, yet how well they handle scientific code remains an open question. Existing datasets and benchmarks are limited in scale, domain coverage, or executable verification, leaving the true gap between current LLMs and reliable scientific code generators inadequately assessed. To address these limitations, we present SciCodePile, the largest scientific code corpus to date, constructed from 37,737 public repositories and collectively comprising 128GB of code that spans multiple computational science disciplines. From this corpus, we further curate an executable benchmark of 200 tasks, each equipped with a sandboxed execution environment and an automated test harness for functional verification. We evaluate 15 LLMs from both open-source and closed-source families on three tasks: prefix-to-suffix completion, fill-in-the-middle infilling, and executable code generation. Results show that scientific code generation remains highly challenging: The best CodeBLEU reaches only 38.13 and 38.37 on the two completion tasks, while the strongest model achieves just 12.30\% Pass@1 on the executable benchmark, underscoring how far current models remain from reliable scientific code generation. To demonstrate the training utility of SciCodePile, we further show that continued pretraining on our corpus improves CodeBLEU by $\times$2.84 on scientific code completion, and instruction tuning on our data improves Pass@1 by $\times$4.79 on the executable benchmark. All code and data are available at https://huggingface.co/SciCodePile.

0 Citations
0 Influential
8.5 Altmetric
42.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!