시계열 기반 모델에서의 예측 성능 저하 현상
Forecast Collapse in Time-Series Foundation Models
본 연구에서는 1,000개의 미국 주식에 대한 시간별 수익률을 예측하는 과정에서 예상치 못한 현상을 관찰했습니다. 바로 예측값이 거의 평탄해지고, 교차 단면 상관관계 측면에서 성능이 저하되는 '예측 성능 저하' 현상입니다. 흥미롭게도 동일한 설정 하에서 거래량을 예측할 때는 이러한 현상이 대부분 사라집니다. 본 연구는 시계열 기반 모델(TSFM), 12개의 심층 학습 예측 모델, 그리고 97개의 공개된 벤치마크 구성을 사용하여 예측 성능 저하 현상을 분석했습니다. 그 결과, 이 현상은 예측 가능성과 밀접한 관련이 있음을 확인했습니다. 예측 성능 저하는 낮은 예측 가능성이 교정된 점 예측의 범위를 제한하고, 개별 시리즈에 대한 최적화가 여러 시리즈 간의 구조를 파악하지 못하기 때문에 발생합니다. 이러한 분석은 교정(calibration)과 순위(ranking) 사이의 균형 문제를 드러냅니다. 제곱 오차를 최소화하는 것은 평탄한 예측을 초래하지만, 교차 단면 상관관계를 직접 최적화하면 순위를 향상시킬 수 있지만 예측 값의 범위를 한 자릿수 이상 증가시킬 수 있습니다. 이러한 상충 관계를 해결하기 위해, 본 연구에서는 교정 및 순위 간의 균형을 맞추는 간단한 목적 함수인 CalibRank를 제안합니다. Finance1K 데이터셋에서 CalibRank는 교차 단면 상관관계를 거의 3배로 향상시키면서 예측 값의 범위를 목표 범위 내에 유지하고, 모든 테스트 모델에서 상관관계를 개선했습니다. 본 연구 결과는 기존 시계열 평가 방법론의 한계를 보여줍니다. 개별 시리즈 지표는 다운스트림 의사 결정에 필요한 여러 시리즈 간의 구조에서의 실패를 숨길 수 있습니다.
When forecasting hourly returns for 1,000 US equities, we observe an unexpected phenomenon: predictions become nearly flat and show poor stock ranking, as measured by cross-sectional correlation. We call this forecast collapse. Surprisingly, the phenomenon largely disappears when forecasting trading volume under the same setting. We investigate forecast collapse across time-series foundation models (TSFMs), twelve deep-learning forecasting models, and 97 public benchmark configurations, and find that it is closely tied to target predictability. We identify two distinct reasons behind it: low predictability limits the amplitude of calibrated point forecasts, while per-series objectives leave cross-series structure unidentified. These findings reveal a calibration-ranking tradeoff: optimizing squared error leads to flat predictions, whereas directly optimizing cross-sectional correlation improves ranking but can inflate forecast amplitude by more than an order of magnitude. To address this tradeoff, we introduce CalibRank, a simple objective that balances calibration and ranking. On Finance1K, CalibRank nearly triples cross-sectional correlation while keeping amplitude close to the target, and improves correlation on all tested models. Our results reveal a blind spot in conventional time-series evaluation: per-series metrics can hide failures in cross-series structure needed by downstream decisions.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.