2605.28258v1 May 27, 2026 cs.SE

지속적인 게임 생성을 위한 GUI 에이전트

GUI Agents for Continual Game Generation

Yuanzhe Shen
Yuanzhe Shen
Citations: 40
h-index: 3
Ruihan Yang
Ruihan Yang
Citations: 604
h-index: 8
Bo Li
Bo Li
Citations: 198
h-index: 3
Qingyi Si
Qingyi Si
Citations: 12
h-index: 1
Hongcheng Guo
Hongcheng Guo
Citations: 12
h-index: 1
Haonan Ge
Haonan Ge
Citations: 41
h-index: 4
Yixu Huang
Yixu Huang
Citations: 2
h-index: 1
Na Li
Na Li
Citations: 332
h-index: 11
Zhe Wang
Zhe Wang
Citations: 127
h-index: 2
Kai Chen
Kai Chen
Citations: 33
h-index: 3
Guangjin Wang
Guangjin Wang
Citations: 25
h-index: 2

게임을 생성하는 것과 플레이 가능한 게임을 만드는 것은 다릅니다. 코드 생성 분야의 발전에도 불구하고, 기존 접근 방식은 게임 생성을 프롬프트에서 결과물로의 일회성 변환으로 취급하여 상호 작용 수준에서의 실패를 감지하지 못합니다. 우리는 게임 생성을 평가하고 개선하는 데 플레이어가 필요하며, 그래픽 사용자 인터페이스(GUI) 에이전트가 이 과정에서 두 가지 중요한 역할을 수행한다고 주장합니다: (1) 객관적인 평가자로서, PlaytestArena라는 새로운 평가 환경을 소개합니다. 이 환경은 웹 브라우저 기반의 200개 게임 생성 작업을 8가지 장르로 분류하고, 예상되는 플레이 행동에 대한 기준을 정의하며, GUI 에이전트가 각 빌드를 브라우저에서 로드하여 실행하면서 이를 평가합니다. (2) 주관적인 테스트 플레이어로서, Play2Code를 제안합니다. 여기서 게임 에이전트와 GUI 에이전트는 공유 메모리를 사용하여 지속적으로 상호 작용하며, 게임 생성을 코딩과 플레이 간의 대화로 변환합니다. 우리의 실험 결과는 최첨단 모델조차도 직접적으로 플레이 가능한 게임을 생성하는 데 어려움을 겪는다는 것을 보여줍니다. 반면, Play2Code는 66.8%의 기준 통과율을 달성했으며, 이는 단일 패스 방식 및 에이전트 기반 코딩의 기본 성능보다 각각 37.1점과 14.6점 향상된 수치입니다. 추가 분석 결과, GUI 테스트 플레이어의 피드백은 인간 보고서보다 추적하기 쉽지만, 인간 테스터와 유사한 고유한 특징을 가지고 있으며, 이는 게임 플레이 테스트를 대화형 코드 생성에 대한 중요한 시험대로 확립합니다. 저희 프로젝트 웹사이트는 https://continual-game-generation.vercel.app/ 에서 확인할 수 있습니다.

Original Abstract

Generating a game is not the same as making one that can be played. Despite advances in code generation, existing approaches treat game generation as one-shot translation from prompt to artifact, leaving interaction-level failures undetected. We argue that evaluating and improving game generation requires a player, and study two roles for graphical user interface (GUI) agents in this process: (1) as an objective evaluator, for which we introduce PlaytestArena, a new evaluation environment that pairs 200 browser-based game generation tasks across eight genres with rubrics of expected in-play behaviors, adjudicated by a GUI agent that loads each build in a browser and plays it; and (2) as a subjective playtester, for which we propose Play2Code, where a game agent and a GUI agent operate in a sustained loop with shared memory, turning game generation into a dialogue between coding and playing. Our experiments show that even frontier models struggle to generate playable games directly, while Play2Code achieves a 66.8\% rubric pass-rate, improving over single-pass and agentic-coding baselines by 37.1 and 14.6 points respectively. Further analysis shows that GUI playtester feedback is more traceable than a human report, yet idiosyncratic in ways reminiscent of human testers, establishing game playtesting as a critical testbed for interactive code generation. Our project website is available at https://continual-game-generation.vercel.app/.

9 Citations
1 Influential
5.5 Altmetric
38.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!