2606.16173v1 Jun 15, 2026 cs.AI

TimeVista: 시계열 예측 모델을 평가하기 위한 비전-언어 모델 활용 연구

TimeVista: Exploring and Exploiting Vision-Language Models as Judges for Time Series Forecasting

Jialong Wu
Jialong Wu
Tsinghua University
Citations: 691
h-index: 10
Mingsheng Long
Mingsheng Long
Citations: 115
h-index: 5
Jianmin Wang
Jianmin Wang
Citations: 80
h-index: 5
Haoran Zhang
Haoran Zhang
Citations: 2,850
h-index: 8
Xin Su
Xin Su
Citations: 63
h-index: 4
Yuxuan Wang
Yuxuan Wang
Citations: 788
h-index: 6
Zhi Chen
Zhi Chen
Citations: 136
h-index: 2
Yong Liu
Yong Liu
Citations: 22
h-index: 2

고품질의 시계열 예측은 실제 의사 결정에 매우 중요합니다. 그러나 기존의 점별 지표는 복잡한 시간적 패턴을 제대로 반영하지 못하며, 인간의 직관적인 선호도와 일치하지 않는 경우가 많습니다. 'LLM-as-a-Judge' 패러다임은 텍스트 평가 방식을 혁신적으로 변화시키며, 유연하고 인간과 유사한 판단을 제공하지만, 이 방법론이 시계열 데이터에 적용된 사례는 아직 미미합니다. 본 연구에서는 비전-언어 모델(VLMs)을 시계열 예측 모델의 평가 기준으로 활용하여, 텍스트 정보를 기반으로 시계열 그래프를 이해하는 VLM의 능력을 활용합니다. 특히, 문맥 정보를 바탕으로 한 미시적 및 거시적 판단을 통합하는 새로운 프레임워크를 제안하여 시계열 예측을 평가합니다. 이를 위해, 우리는 상세한 평가 기준이 포함된 5563개의 시계열 샘플로 구성된 종합적인 VLM-as-a-Judge 벤치마크인 TimeVista를 소개합니다. 광범위한 메타 평가 결과, VLMs는 매우 신뢰할 수 있는 평가 도구이며, 기존 지표보다 인간의 선호도와 훨씬 높은 일관성을 보이는 것으로 나타났습니다. 본 벤치마크를 기반으로, 최근 개발된 시계열 기초 모델(TSFMs)을 VLM-as-a-Judge 패러다임을 통해 종합적으로 평가했습니다. 연구 결과는 VLMs가 견고하고 해석 가능한 평가 기준으로 작용하며, 시계열 모델을 평가하기 위한 포괄적이고 인간 중심적인 기준을 제공한다는 것을 보여줍니다.

Original Abstract

High-quality time series forecasting is pivotal for real-world decision-making. However, traditional point-wise metrics often fail to reveal complex temporal patterns and align poorly with human intuitive preferences. While the ''LLM-as-a-Judge'' paradigm has revolutionized text evaluation by providing flexible, human-aligned judgment, its application to time series remains largely unexplored. In this paper, we leverage Vision-Language Models (VLMs) as judges for time series forecasting, harnessing their ability to comprehend time series plots grounded in textual information. Specifically, we propose a novel framework integrating micro- and macro-level judgments informed by contextual information to evaluate time series forecasting. To this end, we introduce TimeVista, a comprehensive VLM-as-a-Judge benchmark comprising 5563 time series samples paired with detailed evaluation rubrics. Extensive meta-evaluations demonstrate that VLMs are highly reliable judges, achieving significantly higher consistency with human preferences than conventional metrics. Building upon our benchmark, we comprehensively assess recent Time Series Foundation Models (TSFMs) under the VLM-as-a-Judge paradigm. Our results demonstrate that VLMs serve as robust and interpretable judges, providing a comprehensive, human-aligned standard for evaluating time series models.

0 Citations
0 Influential
5 Altmetric
25.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!