2605.27904v1 May 27, 2026 cs.AI

Dr-CiK: 미래 예측 기반 에이전트를 위한 테스트 환경

Dr-CiK: A Testbed for Foresight-Driven Agents

Yihong Tang
Yihong Tang
Citations: 110
h-index: 3
A. Williams
A. Williams
Citations: 403
h-index: 4
Arjun Ashok
Arjun Ashok
Citations: 429
h-index: 5
V. Zheng
V. Zheng
Citations: 72
h-index: 5
Lijun Sun
Lijun Sun
Citations: 19
h-index: 2
Alexandre Drouin
Alexandre Drouin
Citations: 72
h-index: 5
I. Laradji
I. Laradji
Citations: 4,198
h-index: 33
V. Zantedeschi
V. Zantedeschi
Citations: 306
h-index: 8
Étienne Marcotte
Étienne Marcotte
Citations: 42
h-index: 2

실제 환경에서의 시계열 예측은 과거 관측 데이터뿐만 아니라, 노이즈가 많고 다양한 정보 소스에서 적극적으로 탐색해야 하는 외부 맥락에 의존하는 경우가 많습니다. 그러나 기존의 맥락 기반 예측 벤치마크는 일반적으로 필요한 맥락이 이미 제공된다고 가정하며, 에이전트가 스스로 이러한 맥락을 식별할 수 있는지 여부에 대한 질문에는 답하지 않습니다. 따라서 본 논문에서는 에이전트가 문서 데이터베이스에서 예측에 관련된 유용한 맥락을 검색하고, 불필요한 정보를 제거하며, 검색된 맥락을 예측에 활용할 수 있는 증거로 추출하고, 이를 기반으로 예측을 수행하는 능력을 평가하기 위한 벤치마크인 Dr-CiK를 소개합니다. 컨텍스트 제거 실험과 최첨단 심층 학습 및 예측 방법의 결합을 통한 평가를 통해, 고품질 맥락이 Dr-CiK에서 예측 성능을 크게 향상시키는 것을 확인했습니다. 그러나 대부분의 기존 DR 에이전트는 실제 맥락 증거의 극히 일부(대부분 5% 미만)만을 복구하며, 종종 불필요한 정보에 의해 오도당하고 (>80%의 불필요한 정보 인용), 검색된 맥락을 사용할 때 예측 성능이 오히려 저하되는 경우도 발생합니다. 이러한 결과는 미래를 예측하기 위해 적절한 맥락을 탐색하는 미래 예측 기반 에이전트에 대한 연구를 촉진합니다.

Original Abstract

Time series forecasting in real-world settings often depends not only on historical observations, but also on external context that must be actively discovered from noisy, heterogeneous information sources. Yet existing context-aided forecasting benchmarks typically assume that the supporting context is already provided, leaving open whether agents can identify it on their own. Therefore, we introduce Dr-CiK, a benchmark for evaluating whether agents can retrieve forecasting-relevant supporting context from a document corpus, filter out distractors, distill the retrieved context into forecast-useful evidence, and generate forecasts supported by that evidence. Through context ablations and evaluations of state-of-the-art deep research and forecasting methods paired together, we show that high-quality context substantially improves forecasting performance in Dr-CiK. However, most existing DR agents recover only a small fraction of the ground-truth supporting evidence (usually <5%), are frequently misled by distractors (>80% distractor citations), and can cause forecasters to perform worse with retrieved context than without context. Our results motivate research on foresight-driven agents that search for the right context to predict the future.

0 Citations
0 Influential
16.5 Altmetric
82.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!