체스 해설에 대한 환각 현상: 도구 기반 평가를 통한 LLM 성능 분석
Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary
체스와 같은 분야에서 초인공지능 수준의 게임 엔진은 전문가 수준의 분석을 쉽게 제공하지만, 이러한 전문 지식을 교육적으로 유용하게 만들기 위해서는 자연어 설명을 제공해야 합니다. 대규모 언어 모델(LLM)은 이 격차를 해소할 수 있지만, 제한된 도메인별 지식으로 인해 종종 환각 현상을 일으키며, 기존의 참조 기반 평가 또는 LLM-as-a-judge 프레임워크로는 이러한 오류를 안정적으로 감지하기 어렵습니다. 본 연구에서는 체스 해설을 기본적인 주장에 분해하고, 엔진 지원 도구와 전문가가 검증한 표준 데이터(gold reference)를 활용하여 사실 정확성, 개념적 범위 및 수 평가를 측정하는 ACT-Eval이라는 평가 프레임워크를 제안합니다. 교육적, 토너먼트, 그리고 중요한 상황을 포함하는 325개의 체스 포지션-수 쌍으로 구성된 데이터셋을 공개하며, 여기에는 전문가가 검증한 표준 데이터 125개와 5가지 오류 분류 체계가 포함되어 있습니다. 주요 상용 및 오픈 소스 모델을 평가한 결과, 사실 기반 환각 현상은 여전히 체스 해설에서 광범위하게 나타나는 것으로 확인되었습니다. 도구 사용 없이 GPT-5.4는 22.0%의 부정확한 하위 주장을 생성하는 반면, 더 작은 오픈 소스 모델은 40%를 초과합니다. 도구를 활용하면 사실 정확성과 수 평가가 크게 향상되지만, 모든 모델에서 전문가 수준의 전략적, 전술적 아이디어에 대한 이해는 여전히 제한적인 것으로 나타났습니다. 인간 검증 결과, ACT-Eval의 사실 판단은 관찰된 인간 합의 범위 내에 있으며, 커버리지 점수는 인간이 평가한 전략적 완전성과 높은 상관관계를 보입니다.
Superhuman game engines in domains like chess have made expert-level evaluations easily accessible, yet they communicate what is true without the natural-language explanations that make such expertise educationally useful to experts and non-experts alike. Large language models could, in principle, bridge this gap, but they frequently hallucinate due to limited domain-specific knowledge, and standard reference-based or LLM-as-a-judge frameworks cannot reliably detect these errors. In this work, we present ACT-Eval, an evaluation framework that decomposes chess commentary into atomic claims and routes them to engine-supported tools and expert-annotated gold references to assess factual correctness, conceptual coverage, and move-quality judgment. We release a benchmark of 325 position--move pairs spanning pedagogical, tournament, and critical positions, including 125 positions with expert-verified gold atoms and a five-class error taxonomy. Evaluating leading proprietary and open-weight models, we find that factual hallucinations remain pervasive in chess commentary: GPT-5.4 without tools produces incorrect sub-claims 22.0% of the time, while smaller open-weight models exceed 40%. Although tool augmentation substantially improves factual correctness and move-quality assessment, coverage of expert strategic and tactical ideas remains limited across all models. Human calibration shows that ACT-Eval's factual judgments fall within the observed range of inter-human agreement, while its coverage scores correlate strongly with human assessments of strategic completeness.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.