2608.03420v1 Aug 04, 2026 cs.AI

경험 메모리를 활용한 LLM 에이전트의 순차적 의사 결정 개선 연구

Towards Improving Sequential Decision-Making in LLM Agents via Experience Memory

Department of Rehabilitation Science
Department of Rehabilitation Science
Citations: 5
h-index: 1
Jakub Rada
Jakub Rada
Citations: 7
h-index: 1
Viliam Lisý AI Center
Viliam Lisý AI Center
Citations: 0
h-index: 0
Faculty of Electrical Engineering
Faculty of Electrical Engineering
Citations: 1,186
h-index: 19
C. Prague
C. Prague
Citations: 499
h-index: 10

대규모 언어 모델(LLM)은 단일 단계 추론 작업에서 상당한 발전을 이루었지만, 순차적 의사 결정 능력에 대한 이해는 아직 부족합니다. 본 연구에서는 완전 관측 환경의 2인 제로섬 게임을 통해 이를 분석합니다. 이러한 게임은 정확한 평가를 제공하며, 규칙에 따라 결과가 결정되고 각 단계의 최적 행동을 계산하거나 근사할 수 있습니다 (판정 모델 불필요). 다양한 수준의 LLM을 테스트한 결과, LLM은 간단한 게임인 티켓택토 또는 커넥트 포에서 최적이 아닌 플레이를 보이며, MCTS(Monte Carlo Tree Search) 알고리즘 기반의 상대방에게 패배하는 경향이 나타났습니다. 게임 트리는 유지하면서 표면 형태만 변경하는 경우에도 성능 변화가 미미한 것으로 나타났는데, 이는 LLM의 성능 저하가 단순히 기억된 전략의 회복 문제로 완전히 설명될 수 없음을 시사합니다. 이러한 성능 격차를 해결하기 위해, 본 연구에서는 순차적 환경에 적합하도록 설계된 경험 메모리를 활용하는 에이전트 프레임워크를 제안하고, 순차적 의사 결정에서 흔히 발생하는 문제인 신용 할당 문제를 해결하고자 합니다. 실험 결과, 게임 종료 후의 분석 및 규칙 추출을 통해 티켓택토 게임에서 LLM의 성능을 향상시킬 수 있으며, 이는 모델 가중치를 수정하지 않고도 달성할 수 있음을 확인했습니다.

Original Abstract

Large language models have improved substantially on single-shot reasoning tasks, but their performance in sequential decision-making is less well understood. We study this on fully-observable two-player zero-sum games, which provide ground-truth evaluation: outcomes are determined by the rules, and optimality of individual moves can be computed or approximated, without relying on a judge model. Across model tiers, LLMs play suboptimally in simple games such as tic-tac-toe or Connect Four, and lose to MCTS opponents. Obfuscations that preserve the game tree but rewrite its surface form leave performance largely unchanged, indicating the gap is not fully explained by recall of memorized strategies. Motivated by this performance gap, we introduce an agentic framework enhanced with an experience memory designed for the sequential setting and addressing common challenges of sequential decision-making such as credit assignment. We show that post-game reflection and rule extraction yield measurable improvements on tic-tac-toe without modifying the model weights.

0 Citations
0 Influential
9.5 Altmetric
47.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!