ForecastBench-Sim: 시뮬레이션 환경 기반 예측 성능 평가 도구
ForecastBench-Sim: A Simulated-World Forecasting Benchmark
일반적인 AI 시스템의 예측 성능을 평가하는 기존 방법은 실제 세계의 제약을 따르는 경우가 많습니다. 즉, 결과가 느리게 나타나고, 극단적인 사건이 드물며, 가설 검증 질문에 대한 점수 부여가 어렵다는 단점이 있습니다. 본 논문에서는 Freeciv라는 턴 기반 전략 게임(Civilization 시리즈를 모방)의 시뮬레이션을 기반으로 구축된 예측 성능 평가 도구인 ForecastBench-Sim을 소개합니다. 예측 모델은 고정된 세계 보고서(현재 게임 상태의 구조화된 스냅샷)를 입력받아 숨겨진 미래 상태에 대한 질문에 답변하고, 벤치마크 시스템은 시뮬레이션을 계속 진행하면서 예측 결과를 평가합니다. 시뮬레이션 환경 덕분에 동일한 설정을 사용하여 임의의 시간 간격으로 연속적이거나 이진 예측 질문을 생성할 수 있으며, 조건부 또는 인과 관계 질문을 위한 개입 환경(intervention worlds)을 구성하고, 희귀하거나 파괴적인 결과에 대한 해결된 예시를 제공할 수 있습니다. 본 논문에서는 벤치마크 파이프라인, 질문 유형, 평가 프로토콜, 관련 자료를 설명하고, 모델 평가 및 익명화된 사용자 시험 결과를 제시합니다. ForecastBench-Sim은 실제 세계 예측 벤치마크를 보완하여 동적인 환경 상태에서 확률적 추론을 연구하기 위한 통제되고 즉시 해결 가능한 과제를 제공하는 것을 목표로 합니다.
Forecasting benchmarks for general-purpose AI systems usually inherit the constraints of the real world: outcomes resolve slowly, tail events are rare, and counterfactual questions are difficult to score. We introduce ForecastBench-Sim, a simulated-world forecasting benchmark built on game rollouts from Freeciv, a turn-based strategy game modelled on the Civilization series. Forecasters receive a fixed world report (a structured snapshot of the current game state) and answer questions about hidden future states; the benchmark then continues the simulation and scores forecasts. Because the world is simulated, the same setup can generate continuous or binary forecasting questions at arbitrary time horizons, paired intervention worlds for conditional or causal questions, and resolved examples of rare or disruptive outcomes. We describe the benchmark pipeline, question families, scoring protocol, and release artifacts, and report validation slices from model evaluations and an anonymized human pilot. ForecastBench-Sim is intended to complement real-world forecasting benchmarks by providing controlled, immediately resolvable tasks for studying probabilistic reasoning under dynamic world states.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.