2605.28751v1 May 27, 2026 cs.LG

외삽적 가중 평균: 코드 강화 학습에서의 정확성-효율성 경계 탐색

Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL

Gabriel Synnaeve
Gabriel Synnaeve
Citations: 63,503
h-index: 58
Jonas Gehring
Jonas Gehring
Citations: 521
h-index: 8
Taco Cohen
Taco Cohen
Citations: 286
h-index: 6
Kunhao Zheng
Kunhao Zheng
Citations: 131
h-index: 5
Pierre Chambon
Pierre Chambon
Citations: 63
h-index: 2
Juliette Decugis
Juliette Decugis
Citations: 135
h-index: 3
Benjamin Négrevergne
Benjamin Négrevergne
Citations: 706
h-index: 15

미세 조정된 체크포인트 간의 선형 보간은 경쟁적인 목표 사이의 파레토 최적점을 추적하는 것으로 나타났지만, 추가적인 강화 학습 없이도 추론 시간에 유용한 새로운 체크포인트를 위해 외삽적 가중 평균이 그러한 경계를 확장할 수 있는지 여부는 불분명합니다. 본 연구에서는 시간 및 메모리 제한 하에서 숨겨진 단위 테스트를 통해 기능적 정확성과 계산 효율성을 동시에 요구하는 경쟁 프로그래밍 분야의 강화 학습 문제를 중심으로 이 질문을 탐구합니다. 공유 초기값에서 시작하여, 우리는 중첩된 단위 테스트 커버리지를 기반으로 체크포인트를 훈련했습니다. 낮은 커버리지 보상은 작은 입력 테스트를 통과해야 하는 반면, 높은 커버리지 보상은 점진적으로 더 큰 테스트를 통과해야 하며, 궁극적으로 전체 테스트 스위트를 통과해야 합니다. 이러한 과정을 통해 정확성-효율성 경계가 나타나는 것을 확인했습니다. 어려운 문제의 경우, 높은 커버리지 보상은 최적화 실패를 줄이지만 정확성 실패를 증가시켜 해결률에 거의 영향을 미치지 않습니다. 낮은 및 높은 커버리지 체크포인트 간의 보간은 이러한 경계를 복구하는 반면, 외삽은 훈련된 끝점을 넘어 경계를 확장합니다. 이러한 경계와 그 외삽적 연속성은 순수한 추론, 도구 사용, 그리고 에이전트 기반 코딩을 포함한 세 가지 추론 설정 및 32B 및 7B 모델 크기 모두에서 나타났습니다. 문제 수준에서, 경계를 따라 이동하면 해결되는 문제의 집합이 변경되며, 이는 외삽된 체크포인트를 추론 시간 스케일링에 대한 상호 보완적인 정책으로 활용할 수 있음을 의미합니다. 동일한 샘플 예산 하에서 최상의 단일 체크포인트보다 외삽적 가중 평균을 사용한 앙상블은 LCB/hard 데이터셋에서 커버리지를 넓히고 pass@250 성능을 3.3% 향상시켰습니다. 이러한 결과는 코드 강화 학습에서의 중첩된 단위 테스트 커버리지가 경계를 유도하며, 외삽적 가중 평균이 이를 탐색하고 확장하며 활용할 수 있음을 보여줍니다.

Original Abstract

Linear interpolation between fine-tuned checkpoints has been shown to trace the Pareto front between competing objectives, but whether extrapolative weight averaging can extend such frontiers to new checkpoints useful at inference time, without additional RL training, remains unclear. We study this question in RL for competitive programming, where hidden unit tests under time and memory limits enforce both functional correctness and computational efficiency. Starting from a shared initialization, we train checkpoints under nested unit-test coverage: low-coverage rewards require passing smaller-input tests, while high-coverage rewards require passing progressively larger tests up to the full suite. This sweep reveals the emergence of a correctness-efficiency frontier: on hard problems, higher-coverage reward reduces optimization failures but increases correctness failures, leaving solve rate nearly unchanged. Interpolation between low- and high-coverage checkpoints recovers this frontier, while extrapolation extends it beyond the trained endpoints. Both the frontier and its extrapolative continuation appear across three inference settings, pure reasoning, tool use, and agentic coding, and across two model scales, 32B and 7B. At the problem level, moving along the frontier changes which problems are solved, making extrapolated checkpoints complementary policies in inference-time scaling. Ensembles with extrapolative weight averaging broaden coverage and improve pass@250 on LCB/hard by 3.3% over the best single checkpoint at matched sample budget. These results show that nested unit-test coverage in code RL induces a frontier that extrapolative weight averaging can navigate, extend, and exploit.

0 Citations
0 Influential
29 Altmetric
145.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!