2605.29653v1 May 28, 2026 cs.AI

PTCG-Bench: LLM 에이전트가 포켓몬 트레이딩 카드 게임을 마스터할 수 있을까?

PTCG-Bench: Can LLM Agents Master Pokémon Trading Card Game?

Renhong Huang
Renhong Huang
Citations: 65
h-index: 3
Dongdong Hua
Dongdong Hua
Citations: 1
h-index: 1
Yang Yang
Yang Yang
Citations: 8
h-index: 2
Yifei Sun
Yifei Sun
Citations: 8
h-index: 2
Feng Gao
Feng Gao
Citations: 11
h-index: 2
Chunping Wang
Chunping Wang
Citations: 487
h-index: 12

전략적으로 복잡한 보드 게임의 경우, 인간 플레이어는 몇 번의 플레이를 통해 빠르게 전략을 익힐 수 있습니다. 자율적인 에이전트는 현실적인 상호작용 환경에서 유사한 능력을 필요로 하지만, 기존 에이전트 벤치마크는 이러한 전략적이고 진화하는 의사 결정 시나리오를 충분히 반영하지 못하는 경우가 많습니다. 본 논문에서는 포켓몬 트레이딩 카드 게임(PTCG)을 기반으로 구축된 벤치마크인 PTCG-Bench를 소개합니다. 이 벤치마크는 LLM 에이전트를 다음 두 가지 상호 보완적인 수준에서 평가합니다. (1) 단일 복잡한 환경 내에서의 의사 결정 성능, 그리고 (2) 축적된 경험을 통해 스스로 발전하는 능력. 또한, 에이전트의 성능과 모델 자체의 능력을 혼동하지 않도록 모듈화된 하니스(harness) 분석 방법을 포함했습니다. 실험 결과, LLM 에이전트는 상당한 수준의 게임 플레이 성능을 달성할 수 있지만, 지속적이고 안정적인 자기 진화는 여전히 어려운 과제이며, 성능은 하니스 설계에 민감하게 영향을 받는다는 것을 확인했습니다. PTCG-Bench가 현실적인 상호작용 환경에서 하니스(harness)를 고려하고 스스로 발전하는 에이전트에 대한 미래 연구를 촉진할 수 있기를 바랍니다.

Original Abstract

Given a strategically complex board game, human players can quickly learn to devise strategies after playing a few rounds. Autonomous agents require similar capabilities in realistic interactive environments, yet existing agent benchmarks often fail to fully capture such strategic and evolving decision-making scenarios. We present PTCG-Bench, a benchmark built on the Pok'{e}mon Trading Card Game (PTCG) that evaluates LLM agents at two complementary levels: (1) their decision-making performance within a single complex environment, and (2) their ability to self-evolving through accumulated experience. We further include a modular harness ablation to better interpret agent performance without conflating it with model capability. Our experiments show that, although LLM agents can achieve non-trivial gameplay performance, sustained and stable self-evolution remains challenging, and performance is sensitive to harness design. We hope that PTCG-Bench will facilitate future research on harness-aware and self-evolving agents in realistic interactive environments.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!