2604.16748v1 Apr 17, 2026 cs.CV

TriTS: 다중 모달 관점에서 본 시계열 예측

TriTS: Time Series Forecasting from a Multimodal Perspective

Xiang Ao
Xiang Ao
Citations: 1,383
h-index: 19

시계열 예측은 금융, 에너지, 교통, 기상 등 다양한 중요한 분야에서 핵심적인 역할을 수행합니다. 그러나 장기 시계열 예측(LTSF)은 여전히 어려운 과제이며, 이는 현실 세계의 신호가 복잡하게 얽힌 시간적 특성을 가지고 있으며, 이러한 특성을 순수한 1차원 관점만으로는 완전히 파악하기 어렵기 때문입니다. 이러한 표현상의 한계를 극복하기 위해, 우리는 1차원 시계열 데이터를 시간, 주파수, 그리고 2차원 이미지 공간으로 투영하는 새로운 모달 간 분리 프레임워크인 TriTS를 제안합니다. Vision Transformer(ViT)의 비효율적인 $O(N^2)$ 계산 복잡도를 피하면서, 1차원-2차원 모달 간의 간극을 원활하게 연결하기 위해, 우리는 Period-Aware Reshaping 전략을 도입하고 Visual Mamba(Vim)를 통합했습니다. 이 접근 방식은 전체 시계열 의존성을 글로벌 시각적 텍스처로 효율적으로 모델링하면서도 선형적인 계산 복잡도를 유지합니다. 또한, 주파수 모달리티를 위해 Multi-Resolution Wavelet Mixing (MR-WM) 모듈을 설계하여, 비정상 신호를 추세 및 잡음 구성 요소로 명시적으로 분리하여 정밀한 시간-주파수 지역화를 달성합니다. 마지막으로, 수치적 안정성을 확보하기 위해 시간 영역에 선형 브랜치를 유지합니다. TriTS는 이 세 가지 상호 보완적인 표현을 동적으로 융합하여 다양한 데이터 환경에 효과적으로 적응합니다. 여러 표준 데이터 세트에 대한 광범위한 실험 결과, TriTS는 최첨단(SOTA) 성능을 달성하며, 기존의 시각 기반 예측 모델보다 파라미터 수와 추론 지연 시간을 크게 줄여 근본적으로 우수한 성능을 보임을 입증합니다.

Original Abstract

Time series forecasting plays a pivotal role in critical sectors such as finance, energy, transportation, and meteorology. However, Long-term Time Series Forecasting (LTSF) remains a significant challenge because real-world signals contain highly entangled temporal dynamics that are difficult to fully capture from a purely 1D perspective. To break this representation bottleneck, we propose TriTS, a novel cross-modal disentanglement framework that projects 1D time series into orthogonal time, frequency, and 2D-vision spaces.To seamlessly bridge the 1D-to-2D modality gap without the prohibitive $O(N^2)$ computational overhead of Vision Transformers (ViTs), we introduce a Period-Aware Reshaping strategy and incorporate Visual Mamba (Vim). This approach efficiently models cross-period dependencies as global visual textures while maintaining linear computational complexity. Complementing this, we design a Multi-Resolution Wavelet Mixing (MR-WM) module for the frequency modality, which explicitly decouples non-stationary signals into trend and noise components to achieve fine-grained time-frequency localization. Finally, a streaming linear branch is retained in the time domain to anchor numerical stability. By dynamically fusing these three complementary representations, TriTS effectively adapts to diverse data contexts. Extensive experiments across multiple benchmark datasets demonstrate that TriTS achieves state-of-the-art (SOTA) performance, fundamentally outperforming existing vision-based forecasters by drastically reducing both parameter count and inference latency.

0 Citations
0 Influential
9.5 Altmetric
47.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!