2605.27000v1 May 26, 2026 cs.CL

더 넓은 범위의 탐색: 코드 추론을 위한 조정된 Pass@K 정책 최적화

Cast a Wider Net: Coordinated Pass@K Policy Optimization for Code Reasoning

Tong Che
Tong Che
Citations: 66
h-index: 3
Suman Banerjee
Suman Banerjee
Citations: 107
h-index: 5
Yilong Li
Yilong Li
Citations: 12,817
h-index: 34

코드 생성을 위해 테스트 시간 컴퓨팅 자원을 할당하는 표준적인 방법은 검증기를 사용한 반복 샘플링이며, 이때 pass@$K$가 대표적인 성능 지표로 사용됩니다. 그러나 기존의 정책 클래스는 $K$개의 독립적인 샘플을 단일 답변 분포에서 추출하므로, 시도 과정에서 종종 유사한 추론 경로로 수렴하여 불필요한 연산을 수행하게 됩니다. 이러한 현상은 경쟁 프로그래밍 환경에서 큰 문제이며, 왜냐하면 많은 문제가 여러 개의 서로 다른 알고리즘 전략을 허용하지만 pass@$K$는 단 하나의 올바른 시도만 요구하기 때문입니다. 본 연구에서는 pass@$K$ 생성을 전략에 대한 공동 탐색으로 전환하는 조정된 Pass@K 정책 최적화 (CPPO)를 제안합니다. CPPO에서 플래너는 $K{=}4$개의 대체 고수준 방법을 담은 튜플을 생성하고, 공유 솔버는 각 방법에 대해 하나의 해를 시도합니다. CPPO는 곱셈적인 플래너 보상 $R_{ ext{plan}} = J_ ext{ψ} imes R_{ ext{out}}$을 사용하여 이 공동 정책을 학습하며, 검증기에 의해 확인된 pass@$K$ 성공으로 이어지는 유효한 전략 튜플에만 기여도를 부여합니다. APPS, CodeContests 및 LiveCodeBench-v6 데이터셋에서 CPPO는 direct sampling, planning 기반 모델, 플래너 전용 SFT (Supervised Fine-Tuning) 모델, 그리고 pass@$K$-지향적인 강화 학습 모델보다 pass@$4$ 성능을 향상시켰습니다. 이 모든 비교는 동일한 $K{=}4$의 솔버 시도 횟수 제한 내에서 이루어졌으며, 9개의 모델-벤치마크 조합 중 6개에서 통계적으로 유의미한 성능 향상을 보였습니다. 가장 큰 단일 성능 향상은 Qwen3.5-9B LiveCodeBench-v6 데이터셋에서 $+0.16$으로 나타났습니다 (0.588 $ ightarrow$ 0.748; paired bootstrap, $p < 0.05$).

Original Abstract

Repeated sampling with a verifier is the standard way to allocate test-time compute for code generation, with pass@$K$ as the canonical metric. Yet the standard policy class draws $K$ independent samples from a single answer distribution, so attempts often collapse onto near-duplicate reasoning paths and waste the budget on redundant rollouts. This failure is costly in competitive programming, where many problems admit multiple distinct algorithmic strategies and pass@$K$ requires only one correct attempt. We propose Coordinated Pass@$K$ Policy Optimization (CPPO), which turns pass@$K$ generation into joint exploration over strategies: a planner emits a tuple of $K{=}4$ alternative high-level methods, and a shared solver attempts one solution per method. CPPO trains this joint policy with a multiplicative planner reward, $R_{\mathrm{plan}} = J_ψ\cdot R_{\mathrm{out}}$, assigning credit only to valid strategy tuples that lead to verifier-confirmed pass@$K$ success. Across APPS, CodeContests, and LiveCodeBench-v6, CPPO improves pass@$4$ over direct sampling, planning baselines, planner-only SFT, and pass@$K$-oriented RL under the same $K{=}4$ solver-attempt budget, with statistically significant gains on six of nine model--benchmark cells. The largest single gain is $+0.16$ on Qwen3.5-9B LiveCodeBench-v6 over the strongest baseline, PKPO ($0.588 \rightarrow 0.748$; paired bootstrap, $p < 0.05$).

1 Citations
0 Influential
17 Altmetric
86.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!