2607.18084v1 Jul 20, 2026 cs.AI

WorldCupArena: 축구 예측 분야의 언어 모델 및 심층 연구 에이전트에 대한 세밀한 평가

WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting

Yihong Tang
Yihong Tang
Citations: 157
h-index: 6
Jiayuan Rao
Jiayuan Rao
Citations: 87
h-index: 4
T. Gui
T. Gui
Citations: 0
h-index: 0
Zhaokai Wang
Zhaokai Wang
Citations: 1,389
h-index: 11
Shangzhe Di
Shangzhe Di
Citations: 554
h-index: 9
Dingli Liang
Dingli Liang
Citations: 0
h-index: 0

축구 경기 결과를 예측하기 위해서는 과거 결과뿐만 아니라 변화하는 정보를 활용하고, 정답이 공개되기 전에 명확한 예측을 내려야 합니다. 본 논문에서는 언어 모델 및 심층 연구 에이전트를 위한 동적 벤치마크인 WorldCupArena를 소개합니다. 2026 FIFA 월드컵이 첫 번째 평가 대상으로 사용되었으며, 동일한 프로세스는 향후 리그 및 대회에도 적용될 수 있습니다. 각 경기 전에 모델은 공통 정보 패키지를 받거나 자체적으로 정보를 검색합니다. 그런 다음 모델은 경기 결과, 점수, 예상 선수 및 이벤트, 경기 통계, 그리고 전체 대회의 결과를 예측합니다. 경기가 끝난 후 이러한 예측은 실제 기록된 결과와 비교됩니다. 우리는 예측 정확도, 정확한 점수 예측 정확도, 그리고 예측 점수가 정확하지 않더라도 어느 정도는 인정해 주는 '점수라인' 점수를 보고하며, 다른 예측 작업에 대한 성능 지표도 함께 제시합니다. 104경기 및 13개의 시스템을 분석한 결과, 유사한 경기 결과 예측 정확도를 가진 모델들 간에도 세부적인 예측에서 차이가 명확하게 나타났습니다. 도박 시장 및 일반 팬의 예측 결과를 기준으로 비교했을 때, 가장 뛰어난 성능을 보이는 시스템은 경기 결과 및 정확한 점수 예측 정확도에서는 미미한 개선을 보였지만, '점수라인' 예측에서는 더 큰 향상을 보였습니다. 새로운 경기 일정은 시작될 때마다 추가할 수 있으며, 이를 통해 벤치마크는 이미 알려진 결과를 사용하지 않고 미래의 모델을 평가할 수 있습니다. 코드, 프롬프트, 예측 결과 및 평가 스크립트는 https://github.com/wzk1015/WorldCupArena 에서 공개적으로 이용할 수 있습니다.

Original Abstract

Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear prediction before the answer is available. We present WorldCupArena, a dynamic benchmark for language models and deep-research agents. The 2026 FIFA World Cup is its first evaluation, and the same process can be reused for future leagues and cups. Before each match, a model either receives a common evidence package or searches for information itself. It predicts the result and score, likely players and events, match statistics, and the outcome of the competition. After the match, these predictions are compared with the recorded result. We report result accuracy, exact-score accuracy, and a scoreline score that gives some credit when a predicted score is close but not exact, together with scores for the other prediction tasks. Across 104 matches and 13 systems, models with similar result accuracy differ more clearly on detailed predictions. Compared with betting-market and human-fan baselines, the best system shows only small gains in result and exact-score accuracy, but a clearer gain in Scoreline. New schedules can be added as they begin, allowing the benchmark to evaluate future models without using outcomes that are already known. Code, prompts, predictions, and evaluation scripts are open sourced at https://github.com/wzk1015/WorldCupArena.

1 Citations
1 Influential
40.222194895832 Altmetric
6.9 Score
Original PDF
18

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!