2606.26487v1 Jun 25, 2026 cs.CL

LLM에 숫자 정보를 전달하는 방법: 시계열 예측을 위한 멀티 웨이블릿 숫자 임베딩

Speaking Numbers to LLMs: Multi-Wavelet Number Embeddings for Time Series Forecasting

Muyan Weng
Muyan Weng
Citations: 25
h-index: 3
Defu Cao
Defu Cao
Citations: 2,500
h-index: 14
Yan Liu
Yan Liu
Citations: 325
h-index: 9
Zijie Lei
Zijie Lei
Citations: 20
h-index: 3
Jiao Sun
Jiao Sun
Citations: 3,664
h-index: 4

대규모 언어 모델(LLM)은 다양한 텍스트 정보를 통합하여 상황 인지적 시계열 예측에 유용하지만, 이들의 이산적인 언어 기반 토큰화 및 임베딩 인터페이스는 연속적인 수치 값과 일치하지 않아 종종 수치 순서 및 예측 신뢰성을 저해합니다. 본 논문에서는 TempoWave라는 플러그 앤 플레이 방식으로 사용할 수 있는 시간 웨이블릿 기반 숫자 인터페이스를 제안합니다. TempoWave는 각 스칼라 관측값을 멀티 웨이블릿, 멀티 스케일 계수를 사용하여 구성된 자릿수별 임베딩으로 변환합니다. TempoWave는 표준 토큰 표현을 직접 대체하여 트랜스포머 호환 형식으로 미세한 국소적 변동과 거시적인 전역 구조를 모두 효과적으로 표현하며, LLM 파이프라인 전체에서 정확한 수치 서식, 뚜렷한 자릿수 식별성 및 일반적인 정규화 작업에 대한 견고성을 유지합니다. 다섯 가지 상황 정보가 풍부한 예측 벤치마크 실험 결과, TempoWave는 표준 숫자 토큰화 및 대체 임베딩 인터페이스를 사용하는 LLM 기반 예측 모델보다 일관되게 성능을 향상시켜 새로운 최고 수준의 결과를 달성했습니다. 이러한 결과는 수치 인터페이스가 중요한 병목 현상임을 강조하며, 체계적인 멀티 해상도 임베딩이 LLM의 상황 추론과 정확한 예측 간의 결합을 개선할 수 있음을 시사합니다. 저희 코드는 https://github.com/DC-research/TempoWAVE 에서, 모델은 https://huggingface.co/Melady/TempoWAVE 에서 확인할 수 있습니다.

Original Abstract

Large language models (LLMs) are attractive for context-aware time series forecasting because they can integrate heterogeneous textual signals, yet their discrete, language-oriented tokenization and embedding interfaces are misaligned with continuous numerical values, often harming numerical ordering and forecasting reliability. We propose TempoWave, a plug-and-play temporal wavelet digit interface that maps each scalar observation into digit-wise embeddings constructed from multi-wavelet, multi-scale coefficients. By directly overriding standard token representations, TempoWave seamlessly exposes both fine-grained local fluctuations and macro global structures in a transformer-compatible form, ensuring that precise numerical formatting, distinct digit identity, and robustness to common normalization operations are maintained throughout the LLM pipeline. Experiments across five context-enriched forecasting benchmarks demonstrate that TempoWave consistently improves LLM-based forecasters over standard numeric tokenization and alternative embedding interfaces, achieving a new state-of-the-art. These results highlight the numeric interface as a key bottleneck and suggest that principled multi-resolution embeddings can better couple LLMs' contextual reasoning with precise forecasting. Our code is available at https://github.com/DC-research/TempoWAVE and our model can be accessed at https://huggingface.co/Melady/TempoWAVE.

2 Citations
0 Influential
30.4657359028 Altmetric
11.0 Score
Original PDF
1

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!