2601.18744v1 Jan 26, 2026 cs.AI

TSRBench: 범용 모델을 위한 포괄적인 멀티태스크 멀티모달 시계열 추론 벤치마크

TSRBench: A Comprehensive Multi-task Multi-modal Time Series Reasoning Benchmark for Generalist Models

Fangxu Yu
Fangxu Yu
Citations: 315
h-index: 7
Haoqiang Kang
Haoqiang Kang
Citations: 436
h-index: 10
Hongyu Zhao
Hongyu Zhao
Citations: 165
h-index: 4
Lianhui Qin
Lianhui Qin
Citations: 490
h-index: 11
Furong Huang
Furong Huang
Citations: 109
h-index: 4
Tianyi Zhou
Tianyi Zhou
Citations: 13
h-index: 2
Xing-ming Guo
Xing-ming Guo
Citations: 466
h-index: 8
Bin Hu
Bin Hu
Citations: 173
h-index: 5
Lingzhi Yuan
Lingzhi Yuan
University of Maryland, College Park
Citations: 41
h-index: 4

시계열 데이터는 실제 환경 어디에나 존재하며 에너지 관리부터 교통 통제에 이르는 중요 애플리케이션에 필수적입니다. 따라서 시계열을 추론하는 능력은 범용 모델이 실무 문제를 해결하는 데 있어 핵심적인 역량입니다. 그러나 기존의 범용 모델 벤치마크에서는 이러한 차원이 현저히 결여되어 있습니다. 이러한 간극을 메우기 위해, 우리는 시계열 추론 능력의 전 범위를 엄격하게 테스트하기 위해 설계된 포괄적 멀티모달 벤치마크인 TSRBench를 제안합니다. TSRBench의 특징은 다음과 같습니다. i) 14개 도메인에 걸친 4,125개의 다양한 문제 세트로 구성되어 있으며, 인식(Perception), 추론(Reasoning), 예측(Prediction), 의사결정(Decision-Making)의 4가지 주요 차원으로 분류됩니다. ii) 이 4가지 차원에서 필수 추론 능력(예: 수치 추론)을 평가하는 15개의 작업을 포함합니다. 광범위한 실험을 통해 TSRBench 환경에서 30개 이상의 주요 상용 및 오픈 소스 LLM, VLM, TSLLM을 평가했습니다. 연구 결과, i) 스케일링 법칙은 인식과 추론에는 유효하지만 예측에서는 적용되지 않았고, ii) 강력한 추론 능력이 정확한 맥락 인식 예측을 보장하지 않아 의미적 이해와 수치 예측 간의 괴리가 있음을 확인했으며, iii) 시계열의 텍스트 및 시각적 입력 표현이 상호 보완적임에도 불구하고 현재의 멀티모달 모델들은 이를 효과적으로 결합하여 상호 성능 이득을 얻지 못하고 있습니다. TSRBench는 현존하는 과제를 조명할 뿐만 아니라 범용 모델의 발전을 위한 귀중한 통찰을 제공하는 표준화된 평가 플랫폼을 제공합니다. 코드와 데이터셋은 https://tsrbench.github.io/ 에서 확인할 수 있습니다.

Original Abstract

Time series data is ubiquitous in real-world scenarios and crucial for critical applications ranging from energy management to traffic control. Consequently, the ability to reason over time series is a fundamental skill for generalist models to solve practical problems. However, this dimension is notably absent from existing benchmarks of generalist models. To bridge this gap, we introduce TSRBench, a comprehensive multi-modal benchmark designed to stress-test the full spectrum of time series reasoning capabilities. TSRBench features: i) a diverse set of 4125 problems from 14 domains, and is categorized into 4 major dimensions: Perception, Reasoning, Prediction, and Decision-Making. ii) 15 tasks from the 4 dimensions evaluating essential reasoning capabilities (e.g., numerical reasoning). Through extensive experiments, we evaluated over 30 leading proprietary and open-source LLMs, VLMs, and TSLLMs within TSRBench. Our findings reveal that: i) scaling laws hold for perception and reasoning but break down for prediction; ii) strong reasoning does not guarantee accurate context-aware forecasting, indicating a decoupling between semantic understanding and numerical prediction; and iii) despite the complementary nature of textual and visual represenations of time series as inputs, current multimodal models fail to effectively fuse them for reciprocal performance gains. TSRBench provides a standardized evaluation platform that not only highlights existing challenges but also offers valuable insights to advance generalist models. Our code and dataset are available at https://tsrbench.github.io/.

7 Citations
1 Influential
5.5 Altmetric
36.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!