2607.29135v1 Jul 31, 2026 cs.LG

HERO: 장기 예측 신경 연산자를 위한 역사 기반 강화 롤아웃 학습

HERO: History-Enriched Rollout Training for Long-Horizon Autoregressive Neural Operators

Jiaquan Zhang
Jiaquan Zhang
Citations: 59
h-index: 4
Yang Yang
Yang Yang
Citations: 83
h-index: 5
Fan Mo
Fan Mo
Citations: 0
h-index: 0
Chaoning Zhang
Chaoning Zhang
Citations: 145
h-index: 7
Shuxu Chen
Shuxu Chen
Citations: 53
h-index: 4
Wei Dong
Wei Dong
Citations: 131
h-index: 6
Yi Lu
Yi Lu
Citations: 3
h-index: 1
Haifan Meng
Haifan Meng
Citations: 0
h-index: 0
Zhihan Lyu
Zhihan Lyu
Citations: 0
h-index: 0

신경 연산자는 시간 의존 미분 방정식(PDE)에 대한 빠른 근사 모델을 제공하며, 학습된 진화 연산자를 자체 예측 결과에 반복적으로 적용합니다. 그러나 이러한 자기 회귀 방식의 롤아웃은 모든 예측 오류를 입력으로 다시 사용하므로, 국부적인 오차가 누적될 수 있습니다. 기존의 롤아웃 학습 전략은 학습 입력과 모델이 생성한 상태 간의 불일치를 줄이는 데 도움이 되지만, 여전히 감독 신호는 실제 경로와 절대적인 차이를 측정하는 데만 의존합니다. 이러한 감독 신호는 모델이 최적화 과정 초기에 나타냈던 장기 예측 실패 문제를 해결했는지에 대한 정보를 제공하지 못합니다. 본 논문에서는 모델의 최적화 이력을 활용하여 상대적인 감독 신호를 추가하는 역사 기반 강화 롤아웃 학습(HERO) 방법을 제안합니다. HERO는 주기적으로 업데이트되는 이전 모델, 현재 모델, 그리고 약간 변형된 입력으로부터 생성된 여러 후보 롤아웃을 롤아웃 오류, 스펙트럼 차이, 에너지 드리프트 및 오차 증가를 기준으로 순위를 매겨 가장 심각한 실패 경로를 참조 대상으로 선택합니다. 이 참조 대상은 마진 기반 객관 함수에 고정된 비교 기준선으로 포함되어 있으며, 이는 독립적인 기울기 방향을 제공하는 대신 실제 롤아웃 기울기에 대한 경계가 있는 샘플 의존적 가중치를 부여합니다. 본 논문에서는 스펙트럼 기반 및 어텐션 기반 구조를 사용한 9개의 PDE 벤치마크에 대한 실험 결과를 통해, HERO가 일관적으로 장기 예측 정확도, 안정적인 롤아웃 길이 및 일반화 성능을 향상시키며 추론 시간 비용은 증가시키지 않는다는 것을 보여줍니다. 이러한 결과는 역사 기반의 상대적인 감독 신호가 장기 자기 회귀 예측을 안정화하는 데 효과적임을 시사합니다.

Original Abstract

Neural operators provide fast surrogates for time-dependent partial differential equations (PDEs) by applying a learned evolution operator recursively to its own predictions, but this autoregressive rollout feeds every prediction error back as input, so local errors accumulate. Existing rollout-training strategies reduce the mismatch between training inputs and self-generated states, yet their supervision still measures only the absolute discrepancy from the ground-truth trajectory. Such supervision is therefore uninformative about whether the operator has overcome the long-horizon failure behaviors it exhibited earlier during optimization. We propose history-enriched rollout training (HERO), which augments conventional absolute trajectory supervision with relative supervision derived from the model's optimization history. HERO ranks detached candidate rollouts from a periodically refreshed lagged operator, the current model, and a perturbed input by rollout error, spectral discrepancy, energy drift, and error growth, and selects the strongest failure trajectory as reference. This reference enters a margin-based objective as a fixed comparison baseline, inducing a bounded, sample-dependent reweighting of the ground-truth rollout gradient rather than an independent gradient direction, which we further analyze theoretically. Experiments on nine PDE benchmarks with spectral and attention-based backbones show that HERO consistently improves long-horizon accuracy, stable rollout length, and out-of-distribution robustness at no inference-time cost. These results indicate that history-enriched relative supervision is effective for stabilizing long-horizon autoregressive prediction.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!