SERPO: 오픈형 테스트 시간 강화 학습을 위한 자체 진화 평가 기준 정책 최적화
SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning
테스트 시간 강화 학습(TTRL)은 언어 모델이 라벨링된 피드백 없이 추론 시간에 스스로 발전할 수 있도록 합니다. 기존 방법들은 답변 투표에 의존하기 때문에, 유효한 응답이 공유된 표준 답변으로 매핑될 수 없는 오픈형 생성 작업에는 자연스럽게 적용하기 어렵습니다. 외부 보상 모델이나 더 강력한 평가자가 없을 경우, 적응은 대신 모델 자체의 출력으로부터 신뢰할 수 있는 보상을 구성해야 합니다. 본 논문에서는 SERPO(Self-Evolving Rubric Policy Optimization)를 소개합니다. SERPO는 답변 투표를 대체하여 응답 증거, 쿼리별 평가 기준 및 정책 매개변수를 공동으로 진화시키는 폐쇄 루프 시스템을 사용합니다. Good-Normal-Bad (G-N-B) 응답 진화는 최대한 분리된 시뮬레이션을 정렬된 아카이브로 구성합니다. 평가 기준 진화는 이러한 아카이브를 구별하는 기준을 유지합니다. 확률적 기준 점수는 판정 토큰의 가능성을 보상 신호로 변환합니다. 정책 진화는 결과 신호를 사용하여 액터(actor)를 최적화합니다. 새로운 액터 시뮬레이션은 아카이브와 평가 기준을 모두 업데이트하여 세 가지 요소의 진화 루프를 완성합니다. 두 가지 모델 구성, 두 가지 동일 도메인 벤치마크 및 네 가지 OOD (Out-of-Distribution, 분포 외부) 벤치마크에서 SERPO는 HealthBench와 ResearchQA에 대해 각각 최대 20.63점과 20.31점을 향상시키고, 여섯 개의 벤치마크의 평균 점수를 최대 8.06점까지 끌어올립니다. 또한 OOD 전이 및 지속적인 교차 벤치마크 진화를 지원합니다.
Test-time reinforcement learning (TTRL) enables language models to self-evolve at inference time without labeled feedback. Existing methods rely on answer voting and therefore do not extend naturally to open-ended generation, where valid responses cannot be mapped to a shared canonical answer. Without external reward models or stronger judges, adaptation must instead construct reliable rewards from the model's own outputs. We introduce SERPO (Self-Evolving Rubric Policy Optimization), which replaces answer voting with a closed loop that co-evolves response evidence, query-specific rubrics, and policy parameters. Good-Normal-Bad (G-N-B) response evolution organizes maximally separated rollouts into ordered archives; rubric evolution retains criteria that discriminate these archives; probabilistic criterion scoring converts verdict-token likelihoods into reward signals; and policy evolution optimizes the actor with the resulting signals. New actor rollouts then refresh both the archives and rubrics, closing the three-way evolution loop. Across two model configurations, two in-domain benchmarks, and four OOD benchmarks, SERPO improves HealthBench and ResearchQA by up to 20.63 and 20.31 points over the corresponding base models, raises the six-benchmark macro-average by up to 8.06 points, and supports OOD transfer and continued cross-benchmark evolution.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.