2607.25308v1 Jul 28, 2026 cs.CL

CAST: 게임 솔버를 활용한 턴 단위 지침 제공을 통한 LLM 에이전트 학습

CAST: Game Solvers as Turn-Level Teachers for LLM Agents

Xunliang Cai
Xunliang Cai
Citations: 74
h-index: 5
Yi-Kai Zhang
Yi-Kai Zhang
Citations: 35
h-index: 3
Yueqing Sun
Yueqing Sun
Citations: 32
h-index: 4
Qi Gu
Qi Gu
Citations: 85
h-index: 5
Ziang Ye
Ziang Ye
Citations: 17
h-index: 3
Yu Wang
Yu Wang
Citations: 58
h-index: 2
Han-Jia Ye
Han-Jia Ye
Citations: 287
h-index: 6
Lan-Zhe Guo
Lan-Zhe Guo
Citations: 66
h-index: 3
Wentao Shi
Wentao Shi
Citations: 495
h-index: 11
Yuchun Miao
Yuchun Miao
Citations: 290
h-index: 5
Fuli Feng
Fuli Feng
Citations: 154
h-index: 5

대규모 언어 모델(LLM)을 장기적인 게임 환경에서 작동하도록 훈련하는 것은 일반적인 의사 결정 능력을 향상시키는 유망한 방법입니다. 그러나 검증 가능한 보상을 사용하는 강화 학습(RLVR)은 성공에 영향을 미치는 결정을 파악하기 어려운 희소한 최종 보상에 의존합니다. 보다 자세한 과정 신호는 이러한 누락된 턴 단위의 기여도를 제공할 수 있지만, 기존 방식으로는 비용 효율성과 정확성을 동시에 유지하기 어렵습니다. 우리는 게임 솔버의 상태 값이 변하는 것이 특정 행동이 성공으로 이어지는 방향으로 상태를 전진시키는지 여부를 나타낸다는 점을 발견했습니다. 이러한 통찰력을 바탕으로, 본 논문에서는 'CAST (Credit Assignment from Solver Teachers)'라는 방법을 제안합니다. CAST는 이러한 상태 값 변화를 솔버의 이점으로 변환하여 RLVR에 턴 단위 신호로 주입하는 방식입니다. 또한, 거의 최적화된 솔버를 가정할 때, 솔버 이점을 극대화하는 것은 솔버로부터 온-폴리시 방식으로 지식을 전달하는 것과 동일하며, 이를 위해 교사 모델의 로짓 값 대신 스칼라 값만 필요하다는 것을 보여줍니다. Sokoban, Minesweeper 및 Rush Hour 게임에서 CAST는 훈련된 모든 기준 모델보다 뛰어난 성능을 보였으며, 도메인 내 환경뿐 아니라 새로운 난이도의 환경에서도 우수한 성능을 나타냈습니다. 또한 ALFWorld 및 WebShop 데이터셋에서 가장 높은 평균 제로샷 성능을 달성했습니다. 본 논문의 코드는 https://github.com/Wloner0809/CAST 에서 확인할 수 있습니다.

Original Abstract

Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement learning with verifiable rewards (RLVR) relies on sparse final rewards that reveal little about which decisions determine success. Denser process signals could supply this missing turn-level credit, but existing sources are hard to keep both cheap and accurate. We observe that changes in a game solver's state value reveal whether an action advances the state toward success. Building on this insight, we propose CAST (Credit Assignment from Solver Teachers), which converts these value changes into solver advantages and injects them into RLVR as turn-level signals. We further show that, under a soft-optimal solver assumption, maximizing the solver advantage is equivalent to on-policy distillation from the solver, requiring only scalar values rather than teacher logits. Across Sokoban, Minesweeper, and Rush Hour, CAST outperforms all trained baselines on every game under both in-domain and unseen-difficulty evaluation and achieves the highest average zero-shot performance on ALFWorld and WebShop. Our code is available at https://github.com/Wloner0809/CAST.

0 Citations
0 Influential
23 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!