2607.24522v1 Jul 27, 2026 cs.LG

FlowCTS: 플로우 모델의 온라인 연속 경로 감독

FlowCTS: On-policy Continuous Trajectory Supervision of Flow Models

Junxia Zhang
Junxia Zhang
Citations: 8
h-index: 2
Jingbo Zhu
Jingbo Zhu
Citations: 515
h-index: 11
Ziming Zhu
Ziming Zhu
Citations: 10
h-index: 2
Yuan Ge
Yuan Ge
Citations: 103
h-index: 5
Hai Zhao
Hai Zhao
Citations: 235
h-index: 6
Xiaoqian Liu
Xiaoqian Liu
Citations: 126
h-index: 6
Chenglong Wang
Chenglong Wang
Citations: 259
h-index: 10
Tong Xiao
Tong Xiao
Citations: 23
h-index: 3
Bei Li
Bei Li
Citations: 140
h-index: 7
Zhengtao Yu
Zhengtao Yu
Citations: 20
h-index: 2
Kaiyang Ye
Kaiyang Ye
Citations: 5
h-index: 1

온라인 증류(OPD)는 대규모 언어 모델의 추가 학습 과정에서 희소 보상 및 노출 편향 문제를 효과적으로 해결하지만, 이를 플로우 모델에 적용하는 연구는 아직 미흡합니다. 이에, 본 논문에서는 동일한 상태에서 시작된 학생 모델의 후속 경로를 서로 연결하는 Flow Continuous Trajectory Supervision (FlowCTS) 방법을 제안합니다. 경로와 속도장의 관계를 활용하여 시간 가중 속도 일치 상위 경계를 도출하고, 이를 감독 단계 수에 따라 실용적인 목표로 이산화합니다. 다중 참조 설정을 통해, 단일 상태 FlowCTS-OPD는 기존의 KL 기반 OPD보다 빠른 수렴 속도를 보입니다. FlowCTS-OPD는 GenEval 지표를 0.90에서 0.93으로, OCR 지표를 0.90에서 0.92로, PickScore 지표를 22.75에서 23.06으로 향상시켰으며, 모든 목표 지표에서 혼합 보상 강화 학습(RL) 기준 모델보다 우수한 성능을 보였습니다. 추가 분석 결과, 기존의 KL 기반 OPD는 보조 SDE 전이 커널로 인해 시간적 감독 불일치가 발생하는 것을 확인했습니다. 온-라인 설정 외에도, FlowCTS는 기존의 SFT (Supervised Fine-Tuning) 방법보다 일관되게 우수한 성능을 나타내며, 특히 OCR에서 두드러집니다. 또한, 감독 단계를 늘리면 더욱 풍부한 경로 정보를 얻을 수 있지만, 최적화 난이도가 증가하는 경향이 있습니다.

Original Abstract

While on-policy distillation (OPD) effectively addresses sparse rewards and exposure bias in large language model post-training, its extension to flow models remains underexplored. To this end, we propose Flow Continuous Trajectory Supervision (FlowCTS), which matches subsequent student and reference trajectories initialized from the same student-visited state. Using the integral relation between trajectories and velocity fields, we derive a temporally weighted velocity-matching upper bound and discretize it into practical objectives parameterized by the number of supervision steps. Under a multi-reference setup, single-state FlowCTS-OPD outperforms vanilla KL-based OPD with faster convergence. FlowCTS-OPD improves GenEval from 0.90 to 0.93, OCR from 0.90 to 0.92, and PickScore from 22.75 to 23.06, while outperforming a mixed-reward RL baseline across all target metrics. Further analysis reveals a clear temporal supervision mismatch in vanilla KL-based OPD arising from its auxiliary SDE transition kernels. Beyond on-policy setting,FlowCTS also consistently outperforms vanilla SFT , particularly on OCR, while increasing supervision steps exhibit a trade-off between richer trajectory information and greater optimization difficulty.

0 Citations
0 Influential
5.5 Altmetric
27.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!