2606.11816v1 Jun 10, 2026 cs.CL

WorldReasoner: 언어 모델 에이전트가 타당한 추론을 통해 미래 사건을 예측하는지 평가

WorldReasoner: Evaluating Whether Language Model Agents Forecast Events with Valid Reasoning

Yizhou Chi
Yizhou Chi
Citations: 75
h-index: 4
Zifeng Ding
Zifeng Ding
Citations: 14
h-index: 2
Eric Chamoun
Eric Chamoun
Citations: 51
h-index: 3
Andreas Vlachos
Andreas Vlachos
Citations: 3
h-index: 1

실제 세계의 사건을 예측하기 위해서는 언어 모델 에이전트가 불완전하고 시간 제약적인 정보로부터 불확실성 속에서 추론해야 합니다. 그러나 에이전트가 실제로 예측하는지 평가하려면 최종 답변의 정확도 그 이상을 고려해야 합니다. 모델은 단순히 암기된 학습 데이터를 회상하거나, 날조된 증거를 인용하거나, 근거 없는 인과 관계 설명을 생성하여 올바른 결과를 낼 수 있습니다. 우리는 시간적으로 타당한 사건 예측을 위한 평가 프레임워크인 WorldReasoner을 제시합니다. 각 작업에서 에이전트는 해결된 예측 질문, 시뮬레이션된 예측 날짜를 받고 해당 날짜 이전에 사용 가능한 증거만 접근할 수 있습니다. 해결 후에는 제출된 확률 값, 인용된 증거, 선택적 인과 사건 그래프가 프레임워크에 의해 점수화됩니다. WorldReasoner은 세 가지 상호 보완적인 측면을 보고합니다. 첫째, 해결된 답변 대비 결과 품질, 둘째, 인용된 출처를 기준으로 한 증거 품질, 셋째, 해결 후의 참고 그래프를 기준으로 한 추론 품질입니다. 이 벤치마크는 에이전트 기반 생성 파이프라인으로 구축되었으며, 예측 질문을 생성하고, 시간 정보를 포함한 증거를 수집하며, 대규모로 참고 그래프를 구축합니다. 이를 통해 14,141개의 기사에서 추출된 8,087개의 사건에 대한 그래프를 포함하는 345개의 해결된 작업이 생성되었습니다. 여섯 가지 제어된 에이전트 환경에서 시간적으로 타당한 정보 검색은 결과 정확도의 가장 중요한 요인이며, 인과 그래프 구축은 주요 사건 복구 능력을 향상시킵니다. 또한 그래프 기반 예측은 주요 사건 및 관련 출처에 더 강력하게 기반하지만, 에이전트는 여전히 신뢰할 수 있는 확률로 변환하기 위해 근거 있는 증거를 활용하는 데 어려움을 겪고 있습니다.

Original Abstract

Forecasting real-world events requires language-model agents to reason under uncertainty from incomplete, time-bounded information. Yet evaluating whether agents genuinely forecast requires more than final-answer accuracy: a model may be correct by recalling memorized training facts, citing fabricated evidence, or producing an unsupported causal story. We present WorldReasoner, an evaluation framework for temporally valid event forecasting. Each task gives an agent a resolved forecasting question, a simulated forecast date, and access only to evidence available before that date; after resolution, the framework scores the submitted probability, cited evidence, and optional causal event graph. WorldReasoner reports three complementary axes: outcome quality against resolved answers, evidence quality over cited sources, and reasoning quality against post-resolution hindsight graphs. The benchmark is built by an agentic construction pipeline that generates forecasting questions, collects time-stamped evidence, and builds hindsight reference graphs at scale, yielding 345 resolved tasks derived from 14,141 articles with graphs covering 8,087 extracted events. Across six controlled agent settings, temporally valid retrieval is the strongest driver of outcome accuracy; causal graph construction improves key-event recovery; and correct graph-enabled forecasts are more strongly grounded in key events and relevant sources, yet agents still struggle to convert grounded evidence into calibrated probabilities.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!