2606.29932v1 Jun 29, 2026 cs.AI

SAGA: 장면 인식 기반, 목표 진화형 에이전트 - CivRealm 전략 계획을 위한 장기적 접근 방식

SAGA: Scene-Aware, Goal-Evolving Agents for Long-Horizon CivRealm Strategy Planning

Yida Wang
Yida Wang
Citations: 94
h-index: 3
Tian Jin
Tian Jin
Citations: 0
h-index: 0
Shuo Chen
Shuo Chen
Citations: 0
h-index: 0
Liuyu Xiang
Liuyu Xiang
Citations: 109
h-index: 6
Zhaofeng He
Zhaofeng He
Citations: 58
h-index: 4
Yingzhuo Liu
Yingzhuo Liu
Citations: 33
h-index: 3
Yexin Li
Yexin Li
Citations: 50
h-index: 3
P. Li
P. Li
Citations: 0
h-index: 0

복잡한 전략 게임에서의 장기적인 전략 계획은 불완전한 정보와 희소 보상 하에서 여러 의사 결정 영역에 걸친 동시적인 추론을 요구합니다. 기존의 LLM 기반 에이전트는 다음과 같은 세 가지 근본적인 문제점을 가지고 있습니다: (1) 원시 타일 좌표로부터 발생하는 장면 인식 부재, (2) 단일화된 상태 정보 덤프에서 비롯되는 문맥 과부하 및 영역 간 결합, 그리고 (3) 각 에피소드를 독립적으로 취급하는 피상적인 교차 게임 학습. 본 연구에서는 이러한 문제점들을 직접적으로 해결하기 위한 세 가지 메커니즘을 갖춘 LLM 멀티에이전트 프레임워크인 SAGA를 제안합니다: (1) Map-Semantic Scene Graph는 게임 개체 간의 유형화된 공간 관계를 자연어 맥락으로 인코딩하여 전역 토큰 증가 없이 공간 인식 문제를 해결합니다. (2) Tool-Augmented Planner는 필요한 경우 세분화된 영역 상태 정보를 가져오고, 각 영역에 특화된 컨트롤러에게 지시 사항을 전달하여 문맥 과부하, 영역 간 결합 및 기계적 제약 위반을 방지합니다. (3) Dual-Horizon Feedback Loop는 게임 내 목표 생성과 구조화된 교차 게임 인과 관계 분석을 결합하여 수동적인 보상 설계 없이 체계적인 전략 진화를 가능하게 합니다. FreeCiv 환경에서 SAGA는 두 가지 강력한 기준 모델보다 더 낮은 분산으로 가장 높은 평균 문명 점수를 달성했으며, 특히 다중 목표 충돌 상황에서 쉽게 희생되는 인프라 건설 측면에서 모든 기준 모델을 크게 능가했습니다. 또한 SAGA는 대부분의 1대1 게임에서 두 가지 강력한 기준 모델보다 우수한 성능을 보였으며, 출력 토큰 수(주요 디코딩 비용)를 27% 줄였습니다. 교차 게임 진화 모듈을 통해 SAGA는 연속된 다섯 번의 에피소드에서 가장 높은 최종 점수를 달성했습니다. 제거 실험 결과, 각 아키텍처 구성 요소가 이러한 성능 향상에 독립적으로 기여한다는 것을 확인했습니다.

Original Abstract

Long-horizon strategic planning in complex strategy games demands concurrent reasoning across multiple decision domains under imperfect information and sparse reward. Existing LLM-based agents suffer from three systematic failures: scene blindness from raw tile coordinates, context overflow and domain coupling from monolithic state dumps, and shallow cross-game learning that treats each episode in isolation. We present SAGA, an LLM multi-agent framework with three mechanisms each directly targeting one class of failure: (i) a Map-Semantic Scene Graph that encodes typed spatial relations among game entities into per-unit natural-language context, resolving spatial blindness without global token inflation; (ii) a Tool-Augmented Planner that pulls fine-grained domain state on demand and dispatches per-domain directives to dedicated specialist controllers, eliminating context overflow, domain coupling, and mechanical constraint violations; and (iii) a Dual-Horizon Feedback Loop that combines periodic within-game goal generation with structured cross-game causal post-mortem, enabling principled strategic evolution without manual reward engineering. Evaluated on FreeCiv, SAGA attains the highest mean civilization score -- the environment's sole sparse objective reward -- with lower variance than the two strongest baselines, and is the only method that significantly surpasses every baseline on infrastructure construction, the resource axis most readily sacrificed under multi-objective conflict. It outscores the two strongest baselines in most head-to-head games while cutting output tokens (the dominant decoding cost) by 27%. Equipped with the cross-game evolution module, SAGA reaches the highest end-of-chain score across five successive episodes. Ablation studies confirm that each architectural component contributes independently to this advantage.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!