2608.10149v1 Aug 10, 2026 cs.LG

REATS: LLM 기반 추론을 활용한 앙상블 학습으로 적응적인 시계열 예측

REATS: LLM Reasoning-based Ensemble Learning for Adaptive Time Series Forecasting

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Li Zhao
Li Zhao
Citations: 12
h-index: 2
Chang Xu
Chang Xu
Citations: 44
h-index: 3
Hui Sun
Hui Sun
Citations: 50
h-index: 5
Xu Zhang
Xu Zhang
Citations: 38
h-index: 4
Nan Ma
Nan Ma
Citations: 0
h-index: 0
Peng Wang
Peng Wang
Citations: 24
h-index: 2
Wei Wang
Wei Wang
Citations: 39
h-index: 4

실제 세계의 시계열 데이터는 다양하기 때문에, 어떤 단일 예측 모델도 모든 경우에 일관되게 우수한 성능을 보이지 않습니다. 앙상블 학습은 상호 보완적인 모델들의 강점을 결합하여 이러한 문제를 해결하지만, 기존 방법들은 주로 고정된 규칙이나 숫자 입력만을 기반으로 하는 블랙박스 모델에 의존하며, LLM의 추론 능력을 활용하여 해석 가능한 가중치 결정 방식을 제공하지 못합니다. 본 논문에서는 LLM의 추론 능력을 활용하여 텍스트 기반 시계열 패턴 설명과 수치적 특징을 함께 처리하고, 체인-오브-생트(Chain-of-Thought) 추론을 통해 해석 가능하고 샘플에 적응적인 앙상블 가중치를 생성하는 지능형 앙상블 라우터인 REATS를 제안합니다. 효과적인 LLM 기반 앙상블 학습을 위해, 핵심 설계 선택 사항들을 연구하고 다음과 같은 내용을 제시합니다: (i) 원시 시계열 데이터를 고정된 토큰 비용으로 하이브리드 텍스트-수치 표현으로 변환하는 구조화된 입력 파이프라인을 통해, API 의존성 없이 규칙 기반의 체인-오브-생트 구성을 가능하게 하고, 검색된 유사 샘플 정보를 추가합니다; (ii) 다양한 다중 행 가중치 감독 방식을 결합하고 토큰 효율적인 퍼센티지 테이블 형식을 사용하여 수치적 복잡도를 줄이고 LLM의 환각 현상을 완화합니다; (iii) SFT(Supervised Fine-Tuning)와 GRPO(Gradient Reinforcement Policy Optimization)를 결합한 두 단계의 미세 조정 프레임워크를 통해, 연속적인 무한 MSE(Mean Squared Error) 격차를 경계된 신호로 변환하여 정밀도를 높이고, 회귀 기반 GRPO에서 나타나는 균일한 민감도 및 이상치 지배적인 장점 압축 문제를 해결합니다. 8개의 벤치마크 실험 결과는 REATS가 경쟁적인 앙상블 모델보다 우수한 성능을 보이며, 자연어 설명을 제공하고, 새로운 후보 모델에 대한 강력한 전이 학습 및 일반화 능력을 갖춘다는 것을 보여줍니다.

Original Abstract

Due to the diversity of real-world time series, no single forecasting model consistently dominates across all samples. Ensemble learning addresses this by combining complementary model strengths, yet existing methods rely on fixed rules or black-box models based solely on numerical inputs, failing to leverage LLM reasoning for interpretable weighting decisions. We propose REATS, which leverages LLM reasoning capabilities as an intelligent ensemble router that jointly processes textual temporal pattern descriptions and numerical features to produce interpretable, sample-adaptive ensemble weights through chain-of-thought reasoning. To enable effective LLM-based ensembling, we study its key design choices and propose: (i) a structured input pipeline that transforms raw time series into hybrid textual--numerical representations with fixed token cost, enabling rule-based chain-of-thought construction without API dependency, augmented with retrieved similar-sample priors; (ii) a diverse multi-row weight supervision scheme coupled with a token-efficient percentage-table format that reduces numerical complexity and mitigates LLM hallucinations; and (iii) a two-stage fine-tuning framework combining SFT with GRPO, where a reciprocal reward mapping transforms the continuous unbounded MSE gap into bounded signals with amplified near-oracle sensitivity, addressing the uniform sensitivity and outlier-dominated advantage compression inherent in naive reward designs for regression-based GRPO. Experiments on eight benchmarks demonstrate that REATS outperforms competitive ensemble baselines while providing natural language explanations and demonstrating strong transfer learning and out-of-domain generalization to unseen candidate models.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!