2607.26358v1 Jul 29, 2026 cs.LG

탐지 한계 영역에서의 추가 학습: 미세 조정에 대한 게임 이론적 접근

Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning

Nika Haghtalab
Nika Haghtalab
Citations: 4,034
h-index: 26
P. Amortila
P. Amortila
Citations: 297
h-index: 8
Keegan Harris
Keegan Harris
Citations: 89
h-index: 3
Brian Lee
Brian Lee
Citations: 12
h-index: 2
Ian Waudby-Smith
Ian Waudby-Smith
Citations: 590
h-index: 10
Michael I. Jordan
Michael I. Jordan
Citations: 3
h-index: 1

강화 학습(RL) 기반 미세 조정은 언어 모델의 성능을 개선하고 참조 정책으로부터의 편향을 제한하기 위해 널리 사용됩니다. 이러한 균형을 맞추는 일반적인 방법은 KL 정규화를 사용하는 RL 목표이지만, 이 방식 자체로는 정규화 계수를 설정하는 원칙적인 방법을 제공하지 않습니다. 실제로는 이 계수가 일반적으로 휴리스틱하게 선택되거나 하이퍼파라미터 검색을 통해 결정되는데, 이는 학습 비용의 불필요한 증가 또는 바람직하지 않은 보상-유지 균형으로 이어질 수 있습니다. 우리는 이러한 절충점을 명시적인 통계적 해석을 제공하는 게임 이론적 프레임워크를 제안합니다. 구체적으로, 우리는 에이전트가 누적된 보상을 최대화하기 위해 정책을 선택하고, 모니터가 시간이 지남에 따라 정책 출력을 관찰하여 참조 정책으로부터의 편차를 검사하는 순차적인 게임을 연구합니다. 다른 관점에서 시작하지는 않지만, 결과적으로 얻어지는 균형 정책은 최적의 정규화 매개변수에 대한 KL 정규화된 RL 문제의 해로 표현될 수 있으며, 이는 통계적 구별 가능성의 단위당 보상을 최대화하는 것으로 해석할 수 있습니다. 우리는 오목-볼록 분수 프로그래밍 분야의 고전적인 결과를 활용하여, 이 균형 계수를 KL 정규화된 RL 목표로 줄여서 학습하는 원칙적인 방법을 제시함으로써, 표준 미세 조정 파이프라인에 유연하게 통합될 수 있도록 합니다. Qwen3-8B 및 Llama-3.2-1B 모델을 사용한 실험에서, 우리의 방법은 지속적인 학습 환경에서 경쟁력 있는 보상-유지 균형을 제공하며, 우리의 프레임워크가 오픈 소스 모델을 제공하는 API 제공 업체를 감사하는 데 어떻게 사용될 수 있는지 보여줍니다.

Original Abstract

Reinforcement learning (RL) fine-tuning is widely used in language model training to improve model performance on a target task while limiting drift from a reference policy. A standard way to balance this trade-off is via a KL-regularized RL objective, although this formulation does not by itself provide a principled way to set the regularization coefficient. In practice, the coefficient is typically chosen heuristically or via hyperparameter search, which can lead to unnecessary overhead in training cost or undesirable reward-retention trade-offs. We instead propose a game-theoretic framework that gives this trade-off an explicit statistical interpretation. Specifically, we study a sequential game in which an agent chooses a policy to maximize cumulative reward while a monitor observes policy outputs over time and tests for deviations from the reference policy. Although not originating from the same perspective, we show that the resulting equilibrium policy can nonetheless be expressed as the solution to a KL-regularized RL problem for an optimal regularization parameter that can be viewed as maximizing reward per unit of statistical distinguishability. Drawing on classical results from concave-convex fractional programming, we provide a principled method for learning this equilibrium coefficient via reduction to the KL-regularized RL objective, thus allowing for flexible integration into standard fine-tuning pipelines. In experiments with Qwen3-8B and Llama-3.2-1B, we demonstrate that our methods result in competitive reward-retention trade-offs in a continual learning setting, and illustrate how our framework may be used to audit API providers serving open-source models.

0 Citations
0 Influential
13 Altmetric
65.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!