2606.10389v1 Jun 09, 2026 cs.AI

정적 평가를 넘어: 적대적인 게임 환경에서 LLM 기반 전략 진화에 대한 공동 진화 메커니즘

Beyond Static Evaluation: Co-Evolutionary Mechanisms for LLM-Driven Strategy Evolution in Adversarial Games

Jianmin Wu
Jianmin Wu
Citations: 20
h-index: 2
Xiaoming Yuan
Xiaoming Yuan
Citations: 8
h-index: 2
Annan Li
Annan Li
Citations: 48
h-index: 2
Dawei Yin
Dawei Yin
Citations: 18
h-index: 2
Ziyang Zhang
Ziyang Zhang
Harbin Institute of Technology
Citations: 124
h-index: 6
Haoran Li
Haoran Li
Citations: 19
h-index: 2
Z. Ge
Z. Ge
Citations: 15
h-index: 1
Yui Lo
Yui Lo
Citations: 0
h-index: 0
Qian Liu
Qian Liu
Citations: 92
h-index: 4
Bo An
Bo An
Citations: 10
h-index: 1
Dongke Rong
Dongke Rong
Citations: 0
h-index: 0
Jiaqun Liu
Jiaqun Liu
Citations: 0
h-index: 0
Dou Shen
Dou Shen
Citations: 43
h-index: 3

최근 LLM 기반 코드 진화 기술의 발전은 프로그램 생성 및 개선을 반복적으로 수행하여 자동화된 발견을 가능하게 했습니다. 그러나 이러한 방법을 적대적인 다중 에이전트 게임에 적용할 때 근본적인 문제가 발생합니다. 즉, 전략이 향상됨에 따라 평가 환경이 변화하여 고정된 평가기가 신뢰성을 잃고 진화가 정체되는 현상이 나타납니다. 우리는 이 문제를 해결하기 위해 세 가지 메커니즘을 제안합니다. 첫째, 발견된 우수 전략을 상대 에이전트 풀에 통합하는 평가기 공동 진화입니다. 둘째, 노이즈가 많은 소수의 게임 점수를 통계적으로 신뢰할 수 있는 평가로 대체하는 계층적 심층 평가입니다. 셋째, 정체 상태를 극복하기 위해 가장 어려운 상대를 동적으로 가중치를 부여하는 약점 압력입니다. 우리는 이 메커니즘들을 OpenEvolve 및 ShinkaEvolve와 동일한 기초 모델 코드 진화 패러다임을 기반으로 구축된 프레임워크인 FAMOU에 구현했습니다. MCTF 2026 3v3 해상 깃발 탈취 과제에서 FAMOU는 두 가지 핵심 LLM을 사용하여 기준 모델보다 일관되게 높은 성능을 보이며, 최고 종합 점수(0.526)와 새로운 상대 에이전트에 대한 가장 뛰어난 일반화 능력(61.7% 승률)을 달성했습니다. 또한, 각 메커니즘이 성능 향상에 기여한다는 것을 확인하기 위한 실험도 수행되었습니다. 주목할 만한 점은 LLM 변이 과정에서 초기 전략에는 존재하지 않던 전술적 구조가 생성된다는 것입니다. 여기에는 탐색 검색 및 적응형 방어와 같은 기능이 포함됩니다. 이는 코드 수준의 진화가 적대적인 환경에서 비자명한 알고리즘 혁신을 가져올 수 있음을 보여줍니다. FAMOU를 통해 진화된 전략은 AAMAS 2026 MCTF Competition에서 하드웨어 라운드 로빈에서 1위를 차지하고 시뮬레이션에서 3위를 차지하여 실제 적용 가능성을 입증했습니다. 우리의 진화 과정을 통해 개발된 최적화된 구현 및 해당 평가 코드는 다음 위치에서 확인할 수 있습니다: https://github.com/1xiangliu1/FAMOU-CoEvo

Original Abstract

Recent advances in LLM-driven code evolution have enabled automated discovery by iteratively generating and improving programs. However, applying these methods to adversarial multi-agent games introduces a fundamental challenge: the evaluation landscape shifts as strategies improve, causing fixed evaluators to become unreliable and evolution to stagnate. We propose three mechanisms to address this challenge: evaluator co-evolution, which incorporates discovered champions into the opponent pool; hierarchical deep evaluation, which replaces noisy few-game scores with statistically reliable assessments; and weakness pressure, which dynamically up-weights the most difficult opponents to break through plateaus. We implement these mechanisms within FAMOU, a framework built upon the same foundation-model code-evolution paradigm as OpenEvolve and ShinkaEvolve. On the MCTF 2026 3v3 maritime capture-the-flag task, FAMOU consistently outperforms both baselines under two backbone LLMs, achieving the highest combined score (0.526) and the best generalization to unseen opponents (61.7% win rate), while ablations confirm that each mechanism contributes to performance. Notably, the LLM mutation process generates tactical structures entirely absent from the seed strategies -- including lookahead search and adaptive interception -- demonstrating that code-level evolution can produce nontrivial algorithmic innovations in adversarial settings. The FAMOU-evolved strategy further achieved 1st place in the hardware round-robin and 3rd in simulation at the AAMAS 2026 MCTF Competition, validating its real-world transferability. The optimized implementation and corresponding evaluation codes developed through our evolutionary process are available at: https://github.com/1xiangliu1/FAMOU-CoEvo

0 Citations
0 Influential
28.493061443341 Altmetric
0.0 Score
Original PDF
2

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!