2604.05681v1 Apr 07, 2026 cs.AI

LUDOBENCH: LLM의 의사 결정 행동을 평가하기 위한 로도 게임 기반 벤치마크

LUDOBENCH: Evaluating LLM Behavioural Decision-Making Through Spot-Based Board Game Scenarios in Ludo

Dhruv Kumar
Dhruv Kumar
Citations: 2
h-index: 1
O. Jain
O. Jain
Citations: 3
h-index: 1

본 논문에서는 LLM의 전략적 추론 능력을 평가하기 위한 벤치마크인 LudoBench를 소개합니다. LudoBench는 주사위 확률, 조각 획득, 안전 구역 탐색, 홈 경로 진행 등 복잡한 계획 수립을 요구하는 다중 에이전트 보드 게임인 로도 게임을 기반으로 합니다. LudoBench는 12가지의 행동적으로 구별되는 의사 결정 범주에 걸쳐 총 480개의 수작업 시나리오로 구성되며, 각 시나리오는 특정 전략적 선택을 분리하여 제시합니다. 또한, 우리는 랜덤 에이전트, 휴리스틱 에이전트, 게임 이론 에이전트, 그리고 LLM 에이전트를 지원하는 완전한 4인용 로도 게임 시뮬레이터를 제공합니다. 게임 이론 에이전트는 깊이 제한된 탐색을 사용하는 Expectiminimax 검색을 통해, 탐욕적인 휴리스틱을 넘어선 전략적 기준을 제시합니다. 4가지 모델 패밀리에 속하는 6개의 모델을 평가한 결과, 모든 모델이 게임 이론 기준에 40-46% 정도만 일치하는 것으로 나타났습니다. 모델들은 다음과 같은 두 가지 뚜렷한 행동 유형으로 나뉩니다. 첫째는 조각을 완성하지만 개발을 소홀히 하는 '완성자(finisher)' 유형이고, 둘째는 개발은 하지만 완성하지 않는 '건설자(builder)' 유형입니다. 각 유형은 게임 이론 전략의 절반만을 반영합니다. 또한, 동일한 보드 상태에서 과거 이력을 고려한 '앙갚음(grudge)' 프레임을 적용했을 때 모델들이 측정 가능한 행동 변화를 보이는 것으로 나타났으며, 이는 프롬프트 민감성이 중요한 취약점임을 보여줍니다. LudoBench는 불확실성 하에서 LLM의 전략적 추론 능력을 벤치마킹하기 위한 가볍고 해석 가능한 프레임워크를 제공합니다. 모든 코드, 시나리오 데이터셋(480개 항목), 모델 출력 결과는 다음 링크에서 확인할 수 있습니다: https://anonymous.4open.science/r/LudoBench-5CBF/

Original Abstract

We introduce LudoBench, a benchmark for evaluating LLM strategic reasoning in Ludo, a stochastic multi-agent board game whose dice mechanics, piece capture, safe-square navigation, and home-path progression introduce meaningful planning complexity. LudoBench comprises 480 handcrafted spot scenarios across 12 behaviorally distinct decision categories, each isolating a specific strategic choice. We additionally contribute a fully functional 4-player Ludo simulator supporting Random, Heuristic, Game-Theory, and LLM agents. The game-theory agent uses Expectiminimax search with depth-limited lookahead to provide a principled strategic ceiling beyond greedy heuristics. Evaluating six models spanning four model families, we find that all models agree with the game-theory baseline only 40-46% of the time. Models split into distinct behavioral archetypes: finishers that complete pieces but neglect development, and builders that develop but never finish. Each archetype captures only half of the game theory strategy. Models also display measurable behavioral shifts under history-conditioned grudge framing on identical board states, revealing prompt-sensitivity as a key vulnerability. LudoBench provides a lightweight and interpretable framework for benchmarking LLM strategic reasoning under uncertainty. All code, the spot dataset (480 entries) and model outputs are available at https://anonymous.4open.science/r/LudoBench-5CBF/

2 Citations
0 Influential
0.5 Altmetric
4.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!