2607.28022v1 Jul 30, 2026 cs.LG

Flux-OPD: 진화하는 문맥을 활용한 온라인 정책 증류

Flux-OPD: On-Policy Distillation with Evolving Contexts

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Chengzhuo Tong
Chengzhuo Tong
Citations: 48
h-index: 3
Bohan Zeng
Bohan Zeng
Citations: 152
h-index: 8
Yang Shi
Yang Shi
Citations: 69
h-index: 4
Yifan Dai
Yifan Dai
Citations: 15
h-index: 3
Yuanxing Zhang
Yuanxing Zhang
Citations: 951
h-index: 19
Wentao Zhang
Wentao Zhang
Citations: 22
h-index: 2
Wenxuan Liu
Wenxuan Liu
Citations: 572
h-index: 8
Zekun Wang
Zekun Wang
Citations: 59
h-index: 3
Ruixu Zhang
Ruixu Zhang
Citations: 22
h-index: 2
Liu Yang
Liu Yang
Citations: 0
h-index: 0
Bozhou Li
Bozhou Li
Citations: 110
h-index: 4
Daili Hua
Daili Hua
Citations: 19
h-index: 3

개방형 영역에서 대규모 언어 모델 학습은 검증 가능한 보상이 부족하여, 작업 선호도를 효과적인 감독 신호로 명확하게 정의하기 어렵습니다. 문맥 정보는 이러한 선호도를 전달할 수 있지만, 학생 모델에 증류된 후에는 추가적인 감독 정보를 제공하지 못합니다. 따라서 본 연구에서는 학생 모델의 성능에 따라 변화하는 문맥을 활용하고자 합니다. 그러나 진화하는 문맥을 학습 과정에서 직접적인 감독 신호로 사용하는 것은 불안정한 증류 목표와 상충되는 분포를 야기하며, 이를 안정시키고 충돌을 줄이기 위한 메커니즘이 필요합니다. 본 논문에서는 역 KL 목적 함수의 분해를 통해 문맥의 효과를 분석하고, 두 가지 주요 결과를 도출했습니다. 첫째, 학생 모델은 문맥 조건부 교사 모델들의 기하 평균으로 증류됩니다. 둘째, 목적 함수에는 이러한 교사 모델 간의 충돌을 측정하는 충돌 항이 포함되어 있습니다. 이러한 분해 결과를 바탕으로, 본 연구에서는 개방형 영역에서 작업 선호도를 학습하기 위해 진화하는 문맥을 활용하는 OPD(온라인 정책 증류) 패러다임인 Flux-OPD를 제안합니다. Flux-OPD는 문맥 조건부 교사 모델과 문맥 자유 교사 모델 간의 차이를 '문맥 차이 신호'로 처리하고, 이를 문맥 자유 교사 모델에 문맥 수정으로 주입하며, 충돌 항을 사용하여 이러한 수정의 강도를 조절합니다. 개방형 작업에서의 실험 결과는 Flux-OPD가 기존 OPD 패러다임을 능가하는 성능을 보이며, 교사 모델의 감독 신호와 진화하는 문맥을 결합할 수 있는 잠재력을 보여줍니다.

Original Abstract

Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision once distilled into the student, motivating contexts that evolve with student performance. However, directly using evolving contexts as in-training supervision results in an unstable distillation target and conflicting distributions, requiring mechanisms to stabilize target and downweight conflicts. In this paper, we analyze the effect of contexts through a decomposition of the reverse KL objective, revealing two findings: the student is distilled toward the geometric mean of context-conditioned teachers, and the objective contains a conflict term that measures conflicts among these teachers. Based on this decomposition, we propose Flux-OPD, an OPD paradigm that uses evolving contexts as in-training supervision to capture task preferences in open-ended domains. Flux-OPD treats the differences between context-conditioned and context-free teachers as contextual difference signals, injects them as contextual corrections into the context-free teacher anchor, and weights their correction strength using the conflict term as an indicator. Experiments on open-ended tasks show that Flux-OPD outperforms existing OPD paradigms, highlighting the potential to combine teacher supervision with evolving contexts.

0 Citations
0 Influential
9.5 Altmetric
47.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!