2608.08010v1 Aug 08, 2026 cs.LG

강화 학습 후속 훈련을 위한 시간 데이터 기반 모델의 정밀한 이웃 정규화

Ground-Truth Neighborhood Regularization for Reinforcement Learning Post-Training of Time Series Foundation Models

Jianqi Zhang
Jianqi Zhang
Citations: 20
h-index: 1
Wenwen Qiang
Wenwen Qiang
Citations: 569
h-index: 15
Fanjiang Xu
Fanjiang Xu
Citations: 33
h-index: 2
Changwen Zheng
Changwen Zheng
Citations: 92
h-index: 5
Xingyu Zhang
Xingyu Zhang
Citations: 23
h-index: 2
Zeen Song
Zeen Song
Citations: 74
h-index: 4

시계열 예측(TSF)은 다양한 실제 응용 분야에서 중요한 역할을 합니다. 최근, 대규모 데이터셋으로 사전 훈련된 시계열 기반 모델(TSFM)들은 강력한 일반화 능력을 보여주며 TSF의 중요한 패러다임으로 부상했습니다. 그 결과, 강화 학습(RL) 후속 훈련은 하위 작업에서의 성능을 더욱 향상시키는 방법으로서 주목받고 있습니다. 그러나 우리는 특정 예측 영역에서 RL 후속 훈련이 TSFM의 출력 분포를 실제 값과 점차적으로 멀어지게 만들어 성능을 제한할 수 있다는 것을 발견했습니다. 이러한 현상을 extbf{최적 이하의 수렴(suboptimal collapse)}이라고 부릅니다. 우리의 분석에 따르면, 초기 단계에서 실제 값 근처의 고품질 경로를 샘플링하는 데 어려움이 최적 이하의 수렴의 중요한 원인입니다. 이 문제를 해결하기 위해, 우리는 TSFM의 RL 후속 훈련을 위한 정밀한 이웃 정규화(GTN-R) 방법을 제안합니다. GTN-R은 실제 값을 기준으로 고품질 영역을 찾고 모델의 확률 분포를 실제 값 주변으로 이동하도록 유도합니다. 이를 통해 고품질 경로를 샘플링할 가능성을 높이고, 최적 이하의 수렴을 완화하며 성능을 향상시킬 수 있습니다. 또한 GTN-R은 다양한 TSFM RL 방법론에 유연하게 통합될 수 있습니다. 광범위한 실험 결과는 그 효과를 입증합니다.

Original Abstract

Time series forecasting (TSF) plays an important role in a wide range of real-world applications. Recently, time series foundation models (TSFMs), pretrained on large-scale datasets, have demonstrated strong generalization capabilities and emerged as an important paradigm for TSF. Reinforcement learning (RL) post-training has consequently attracted growing attention as a means of further improving their performance on downstream tasks. However, we find that, in certain forecast regions, RL post-training may gradually shift the output distributions of TSFMs away from the ground truth, thereby limiting their performance. We refer to this phenomenon as \textbf{suboptimal collapse}. Our analysis suggests that difficulty in initially sampling high-quality trajectories near the ground truth is an important contributing factor to suboptimal collapse. To address this issue, we propose Ground-Truth Neighborhood Regularization (GTN-R) for RL post-training of TSFMs. GTN-R uses the ground truth as a reference for locating high-quality regions and guides the model's probability mass toward the ground-truth neighborhood. This increases the probability of sampling high-quality trajectories, mitigates suboptimal collapse, and improves performance. Moreover, GTN-R can be flexibly integrated into various RL methods for TSFMs. Extensive experiments show its effectiveness.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!