2608.05030v1 Aug 05, 2026 cs.AI

스코어 행렬에서 축구 상황 인지 시뮬레이션으로: 정확한 스코어 재순위를 위한 감사 가능한 LLM 활용 시스템

From Score Matrices to Football-Aware Match-State Simulation: An Auditable LLM Harness for Exact-Score Reranking

Shaopeng Liang
Shaopeng Liang
Citations: 0
h-index: 0

축구 경기 결과 예측은 강력한 통계적 기반과 어려운 맥락적 이해를 결합해야 하는 과제입니다. 동적인 푸아송 모델은 팀의 강점, 기대 득점 및 일관된 스코어 확률을 추정하지만, 선수 역할, 전술적 대결, 동기 부여 또는 선취골이 경기 흐름에 미치는 영향을 직접적으로 이해하지 못합니다. 대규모 언어 모델(LLM)은 이러한 개념에 대해 추론할 수 있지만, 정확한 확률 계산 엔진으로 작동하지는 않습니다. 본 논문에서는 감사 가능한 정보 활용 시스템을 통해 두 가지 요소를 결합했습니다. 우리는 네 가지 버전을 개발했습니다: V1은 동적인 스코어 기반 딕슨-콜스 기준 모델이며, V2는 LLM의 맥락적 평가를 기대 득점 파라미터로 변환합니다. V3은 스칼라 보정 대신 고정된 스코어 후보 집합에 대한 골별 시뮬레이션을 사용하고, V4는 선취골 및 추가 득점 이후의 연쇄 반응 예측, 시간 인지 중단 기능 및 결정론적인 후반부 시나리오를 추가합니다. 이 시스템은 입력 의미 체계를 정의하고, 경기 전 정보를 제공하며, LLM이 검토 가능한 추론 경로를 따르도록 제한합니다. 2025-26 시즌 잉글랜드 프리미어 리그의 처음 150경기 데이터를 사용하여 V1은 정확한 스코어 예측에서 10.0%의 Top-1 및 26.7%의 Top-3 정확도를 달성했습니다. V3은 각각 12.0%와 30.0%, V4는 각각 14.7%와 30.7%를 기록했습니다. V4는 후보군 범위를 77.3%에서 84.7%로 확장했지만, 추가된 후반부 시나리오 중 어느 것도 Top-3 정확한 예측으로 이어지지는 않았습니다. V1의 기본 승패무(1X2) 분포는 53.3%의 argmax 정확도, 0.9878의 로그 손실, 0.5870의 브리에르 점수 및 0.2095의 순위 확률 점수를 달성했습니다. 이러한 결과는 탐색적인 성격을 가지며, 개발된 모델은 완벽한 기준 모델이 아니며, 시간 기반 입력 격리가 폐쇄형 LLM 내에서 결과 기억을 완전히 배제할 수 없습니다. 본 논문의 기여점은 감사 가능한 하이브리드 아키텍처, 명확한 설계 발전 및 축구 상황 인지 시뮬레이션이 스코어 선택에 어떤 영향을 미치는지 보여주는 부정적인 결과를 포함합니다.

Original Abstract

Football score forecasting combines a strong statistical core with a difficult contextual edge. Dynamic Poisson-family models estimate team strength, expected goals, and coherent score probabilities, but do not directly understand roles, tactical matchups, motivation, or how a first goal changes behaviour. Large language models (LLMs) can reason about such concepts, yet are not calibrated probability engines. We combine both components through an auditable information harness. This paper documents four iterations: V1, a dynamic score-driven Dixon-Coles baseline; V2, which maps LLM contextual ratings back into expected-goal parameters; V3, which replaces scalar correction with goal-by-goal simulations over a frozen score-candidate set; and V4, which adds shared first-breakthrough and post-goal cascade judgments, time-aware stopping, and deterministic tail candidates. The harness defines input semantics, supplies pre-match evidence, and constrains the LLM to an inspectable reasoning route. On a chronological replay of the first 150 matches of the 2025-26 English Premier League, V1 achieved 10.0% Top-1 and 26.7% Top-3 exact-score accuracy. V3 reached 12.0% and 30.0%, while V4 reached 14.7% and 30.7%. V4 increased candidate coverage from 77.3% to 84.7%, although no added tail candidate became a Top-3 exact hit. V1's native 1X2 distribution achieved 53.3% argmax accuracy, 0.9878 log loss, 0.5870 Brier score, and 0.2095 ranked probability score. These results are exploratory: the development slice is not an untouched benchmark, and temporal input isolation cannot exclude outcome memory in a closed LLM. The contribution is an auditable hybrid architecture, a clear design evolution, and negative findings showing where football-aware simulation does and does not improve score selection.

0 Citations
0 Influential
0 Altmetric
0.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!