월드컵 아레나: 실시간 토너먼트에서 최첨단 LLM의 잠재력과 데이터 유출 없는 평가
WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament
대규모 언어 모델(LLM)의 예측 능력을 측정하는 대부분의 벤치마크는 사후적으로 이루어집니다. 즉, 이벤트가 이미 발생했고, 정답은 웹에 존재하며, 평가는 암기 여부를 방지해야 합니다. 본 연구에서는 이러한 접근 방식과 반대되는 설계를 제시합니다. 2026년 FIFA 월드컵의 39일 동안, 확장된 추론 능력과 자체 서버 측 웹 검색 기능을 갖춘 6개의 최첨단 LLM에 대해, 매 경기 시작 전에 모든 경기에 대한 예측 카드를 작성하도록 요청했습니다. 이 예측 카드에는 104경기 결과뿐만 아니라, 각 조별 리그의 우승팀 12개와 전체 토너먼트 우승팀을 예측하는 항목이 포함되었습니다. 질문이 주어졌을 때 정답은 존재하지 않았으므로, 평가는 필터링보다는 설계 자체로 데이터 유출을 방지합니다. 결과적으로 총 4,494개의 예측 결과를 보유한 고정된 아카이브가 생성되었습니다. 본 연구는 6개 시스템이 공유하는 행동 양식을 밝힙니다. 경기 결과 예측의 정확도는 평균 63.9%이며, 이는 도박 회사의 가장 유력한 후보를 선택하는 것과 유사합니다. 시스템들은 옳다고 말하기보다는 서로 동의하는 경향이 더 강하며, 따라서 다수결 투표는 큰 의미가 없습니다. 또한, 무승부와 골 가능성을 과소평가하고, 경기 결과 예측을 특정 패턴으로 쏠리는 경향을 보입니다. 정확도는 경기 간 불균형 정도를 반영하며, 알려진 정보의 양과는 직접적인 관련이 없습니다. 따라서, 가장 접전이 예상되는 경기에서는 정확도가 낮아지는 반면, 전체 토너먼트에 대한 질문에는 비교적 잘 답변합니다. 현재 최첨단 시스템 세대는 이 특정 작업에서 뚜렷한 차이를 보이지 않습니다. 상위 및 하위 순위는 유지되지만 중간 순위는 변동하며, 전반적으로 순위 간의 격차는 매우 좁습니다. 본 연구에서는 경기 정보, 일정, 공식 결과 데이터를 벤치마크 자료로 공개하며, 함께 점수 계산 코드를 제공합니다.
Benchmarks that measure the forecasting ability of large language models are almost always retrospective: the event has happened, the answer is somewhere on the Web, and the evaluation must defend itself against memorisation. We report the opposite design. Over the 39 days of the 2026 FIFA World Cup, six frontier LLMs -- all with extended thinking and native server-side web search -- were asked before every kickoff, one match at a time, to fill in a seven-market prediction card for all 104 matches, plus 12 group winners and a pre-tournament outright pool; no answer existed when the question was asked, so the evaluation is leakage-free by construction rather than by filtering, and the frozen archive holds 4,494 scored predictions. What the tournament establishes is a set of behaviours the six systems share. On match outcome they average 63.9%, level with backing the bookmaker's favourite -- which is in fact what they usually do. They agree with one another far more often than they are right, so a majority vote adds nothing. They under-commit to draws and to goals, and crowd their scoreline picks onto a single prototypical result. Accuracy tracks how lopsided a fixture is rather than how much is known about it: it collapses in the closest ties, where the dossiers are richest, while questions about the tournament as a whole are answered well. On this task the current generation of frontier systems is not sharply differentiated: the standings hold up at the top and the bottom across the run and churn in the middle, and the margins stay narrow throughout. The briefing dossiers, fixtures and official results are released as a benchmark, together with the scoring code.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.