더 넓은 범위의 탐색: 코드 추론을 위한 조정된 Pass@K 정책 최적화
Cast a Wider Net: Coordinated Pass@K Policy Optimization for Code Reasoning
코드 생성을 위해 테스트 시간 컴퓨팅 자원을 할당하는 표준적인 방법은 검증기를 사용한 반복 샘플링이며, 이때 pass@$K$가 대표적인 성능 지표로 사용됩니다. 그러나 기존의 정책 클래스는 $K$개의 독립적인 샘플을 단일 답변 분포에서 추출하므로, 시도 과정에서 종종 유사한 추론 경로로 수렴하여 불필요한 연산을 수행하게 됩니다. 이러한 현상은 경쟁 프로그래밍 환경에서 큰 문제이며, 왜냐하면 많은 문제가 여러 개의 서로 다른 알고리즘 전략을 허용하지만 pass@$K$는 단 하나의 올바른 시도만 요구하기 때문입니다. 본 연구에서는 pass@$K$ 생성을 전략에 대한 공동 탐색으로 전환하는 조정된 Pass@K 정책 최적화 (CPPO)를 제안합니다. CPPO에서 플래너는 $K{=}4$개의 대체 고수준 방법을 담은 튜플을 생성하고, 공유 솔버는 각 방법에 대해 하나의 해를 시도합니다. CPPO는 곱셈적인 플래너 보상 $R_{ ext{plan}} = J_ ext{ψ} imes R_{ ext{out}}$을 사용하여 이 공동 정책을 학습하며, 검증기에 의해 확인된 pass@$K$ 성공으로 이어지는 유효한 전략 튜플에만 기여도를 부여합니다. APPS, CodeContests 및 LiveCodeBench-v6 데이터셋에서 CPPO는 direct sampling, planning 기반 모델, 플래너 전용 SFT (Supervised Fine-Tuning) 모델, 그리고 pass@$K$-지향적인 강화 학습 모델보다 pass@$4$ 성능을 향상시켰습니다. 이 모든 비교는 동일한 $K{=}4$의 솔버 시도 횟수 제한 내에서 이루어졌으며, 9개의 모델-벤치마크 조합 중 6개에서 통계적으로 유의미한 성능 향상을 보였습니다. 가장 큰 단일 성능 향상은 Qwen3.5-9B LiveCodeBench-v6 데이터셋에서 $+0.16$으로 나타났습니다 (0.588 $ ightarrow$ 0.748; paired bootstrap, $p < 0.05$).
Repeated sampling with a verifier is the standard way to allocate test-time compute for code generation, with pass@$K$ as the canonical metric. Yet the standard policy class draws $K$ independent samples from a single answer distribution, so attempts often collapse onto near-duplicate reasoning paths and waste the budget on redundant rollouts. This failure is costly in competitive programming, where many problems admit multiple distinct algorithmic strategies and pass@$K$ requires only one correct attempt. We propose Coordinated Pass@$K$ Policy Optimization (CPPO), which turns pass@$K$ generation into joint exploration over strategies: a planner emits a tuple of $K{=}4$ alternative high-level methods, and a shared solver attempts one solution per method. CPPO trains this joint policy with a multiplicative planner reward, $R_{\mathrm{plan}} = J_ψ\cdot R_{\mathrm{out}}$, assigning credit only to valid strategy tuples that lead to verifier-confirmed pass@$K$ success. Across APPS, CodeContests, and LiveCodeBench-v6, CPPO improves pass@$4$ over direct sampling, planning baselines, planner-only SFT, and pass@$K$-oriented RL under the same $K{=}4$ solver-attempt budget, with statistically significant gains on six of nine model--benchmark cells. The largest single gain is $+0.16$ on Qwen3.5-9B LiveCodeBench-v6 over the strongest baseline, PKPO ($0.588 \rightarrow 0.748$; paired bootstrap, $p < 0.05$).
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.