2608.05238v1 Aug 05, 2026 cs.LG

지각과 설명의 분리: 다변량 시계열 데이터와 언어 간의 연산 기반 표현 정렬

Decoupling Perception from Description: Computation-Grounded Representation Alignment between Multivariate Time Series and Language

Chenxi Liu
Chenxi Liu
Citations: 576
h-index: 8
Xinran Feng
Xinran Feng
Citations: 0
h-index: 0
Yi Xie
Yi Xie
Citations: 19
h-index: 3
Chaolin Zhang
Chaolin Zhang
Citations: 0
h-index: 0
Ruikun Li
Ruikun Li
Tsnghua University
Citations: 69
h-index: 5
Wanyun Ling
Wanyun Ling
Citations: 0
h-index: 0
Ziyue Li
Ziyue Li
Citations: 22
h-index: 2

다중 모드 모델을 훈련하여 시계열 데이터를 언어와 연결하는 과정에서 자가 지도 학습의 함정에 빠질 수 있습니다. 일반적인 방법은 LLM이 시계열 데이터를 읽고 설명을 생성하도록 하는 것인데, 이 경우 라벨 품질은 모델이 학습해야 할 지각 능력에 의해 제한됩니다. 데이터는 라벨러가 이미 알고 있는 것 이상을 가르칠 수 없습니다. 또 다른 문제는 대부분의 데이터셋이 단일 변수를 사용하는 반면, 중요한 패턴(채널 간 상관 관계, 선행/후행 구조, 동시에 발생하는 이상 현상)은 여러 변수에서만 나타나는데, 이는 라벨링 LLM의 한계가 가장 잘 드러나는 부분입니다. 이러한 두 가지 문제는 삼자택일을 야기합니다. 기존 방법들은 신뢰성, 현실성 또는 확장성을 갖지만, 이 세 가지를 모두 달성하는 방법은 없습니다. 우리는 지각과 설명을 분리하여 이 문제를 해결했습니다. 결정론적인 코드는 실제 오픈 소스 다변량 시계열 데이터에서 통계 정보를 계산하고, LLM은 이러한 미리 계산된 사실을 언어로 표현합니다. LLM이 어려워하는 지각 능력은 연산에 의해 처리되고, LLM은 표현 능력을 담당합니다. 이를 통해 CGTime이라는 40억 개의 파라미터를 가진 연산 기반 시계열-언어 모델을 개발했습니다. CGTime은 다변량 이해 작업에서 훨씬 더 큰 범용 모델보다 뛰어난 성능을 보입니다. 우리의 테스트 데이터셋에서 가장 높은 다변량 사실 점수를 달성했으며 (GPT-4o-mini의 0.173과 GPT-5.4-nano의 0.203에 비해 0.283), Holm 수정된 쌍별 유의성 검정을 통해 모든 기준 모델과의 차이가 유지됩니다. 또한 생성된 설명에서 검증 가능한 수치 사실을 더 정확하게 제시하며, 더 넓은 범위의 통계적 특성을 다룹니다.

Original Abstract

Training multimodal models to align time series with language runs into a self-supervision trap. The usual recipe asks an LLM to read a series and write a description, so label quality is capped by the perceptual skill the model is supposed to learn. The data can never teach more than the labeler already knows. A second gap makes this worse: most datasets use a single variable, but the patterns that matter (cross-channel correlation, lead-lag structure, co-occurring anomalies) appear only with several variables, right where the labeling LLM's limits are most exposed. These two problems create a trilemma: existing methods are reliable, realistic, or scalable, but none achieves all three. We resolve this by decoupling perception from description. Deterministic code computes a set of statistics from real, open-source multivariate series; the LLM verbalizes those precomputed facts. Perception, which LLMs do poorly, is handled by computation, while the LLM handles expression. This produces CGTime, our 4B-parameter computation-grounded time-series-language model. CGTime outperforms far larger general-purpose models on multivariate understanding tasks: it attains the best multivariate fact score on our held-out benchmark (0.283 vs. 0.173 for GPT-4o-mini and 0.203 for GPT-5.4-nano), a gap that survives Holm-corrected paired significance tests against every baseline. It also states verifiable numerical facts in generated captions more accurately and covers a broader range of statistical properties.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!