2605.00754v4 May 01, 2026 cs.SE

Themis: 유연한 다중 기준 평가를 위한 강력한 다국어 코드 보상 모델 학습

Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring

Iryna Gurevych
Iryna Gurevych
Citations: 1,356
h-index: 17
Indraneil Paul
Indraneil Paul
Citations: 137
h-index: 4
Goran Glavas
Goran Glavas
Citations: 124
h-index: 6

보상 모델(RM)은 언어 모델(LM)의 후속 학습 과정에서 정책 정렬 및 추론 시간 성능 향상을 가능하게 하는 필수적인 요소로 자리 잡았습니다. 그러나 코드 생성 분야에서의 RM 적용에 대한 연구는 상대적으로 부족하며, 기존 연구는 주로 실행 결과 피드백에 초점을 맞추고 있습니다. 이러한 접근 방식은 후속 학습을 기능적 정확성 최적화에만 국한시켜 완전하고 실행 가능한 코드를 평가하는 데 제약을 둡니다. 본 논문에서는 다국어 및 다중 기준 코드 RM의 학습 및 평가를 연구합니다. 이를 위해, 우리는 다섯 가지 선호도 차원(즉, 기준)과 여덟 가지 프로그래밍 언어를 포괄하는 코드 RM 평가 벤치마크인 Themis-CodeRewardBench를 구축하고, 50개 이상의 코드, 수학 및 범용 RM을 분석했습니다. 현재 RM이 기능적 정확성 평가 외에는 제한적인 성능을 보임을 확인한 후, 우리는 현재까지 가장 큰 오픈 소스 코드 선호도 데이터셋인 Themis-CodePreference(35만 쌍 이상의 선호도 페어)를 개발하고, 이를 활용하여 6억에서 320억 개의 파라미터 크기를 갖는 다국어 코드 보상 모델 모음인 Themis-RM을 학습했습니다. 우리의 실험 결과와 분석은 긍정적인 성능 향상 추세를 보여주며, 다양한 선호도를 사용하여 학습할 때 강력한 교차 언어 전이 효과가 있음을 입증합니다. 또한, 신뢰성 있는 코드 보상 모델링을 위해서는 다중 기준 학습의 중요성을 강조합니다.

Original Abstract

Reward models (RMs) have become an indispensable fixture of the language model (LM) post-training playbook, enabling policy alignment and test-time scaling. Research on the application of RMs in code generation, however, has been comparatively sparse, with existing work largely focusing on execution feedback. This choice constrains post-training to optimizing functional correctness over self-contained executable code. In this work, we examine the training and evaluation of multilingual, multi-criteria code RMs. To this end, we first compile Themis-CodeRewardBench, a benchmark to evaluate code RMs across five preference dimensions (i.e., criteria) and eight programming languages, on which we profile 50+ code, math, and general-purpose RMs. Observing the limited proficiency of current RMs beyond scoring for functional correctness, we develop Themis-CodePreference, the largest open-source collection of code preferences to date (more than 350k preference pairs), and use it to train Themis-RM, a suite of multilingual code reward models for flexible multi-criteria scoring, ranging in size from 600M to 32B parameters. Our experiments and ablations demonstrate positive scaling trends, strong cross-lingual transfer when training on diverse preferences, and the importance of multi-criteria training for reliable code reward modeling.

0 Citations
0 Influential
8.5 Altmetric
42.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!