2606.12281v1 Jun 10, 2026 cs.MA

CCKS: 합의 기반 통신 및 지식 공유

CCKS: Consensus-based Communication and Knowledge Sharing

Xiaowei Lv
Xiaowei Lv
Citations: 5
h-index: 1
Deying Li
Deying Li
Citations: 313
h-index: 10
Jinyuan Zu
Jinyuan Zu
Citations: 0
h-index: 0
Yongcai Wang
Yongcai Wang
Citations: 4
h-index: 1
Yunjun Han
Yunjun Han
Citations: 108
h-index: 6
Wenping Chen
Wenping Chen
Citations: 727
h-index: 15
Fengyi Zhang
Fengyi Zhang
Citations: 36
h-index: 4
Naiqi Wu
Naiqi Wu
Citations: 1
h-index: 1

분산 학습 및 분산 실행(DTDE) 환경에서의 협력적 다중 에이전트 강화 학습(MARL)에서, 행동 조언 기반의 지식 공유는 해석 가능하고 확장 가능한 협력을 촉진합니다. 그러나 현재의 행동 조언 방식은 종종 교사의 지침에 지나치게 의존하며, 교사-학생 간의 적합성을 평가하지 않아 과도한 조언, 최적 이하의 안정성 및 성능 저하를 초래합니다. 이러한 문제점을 해결하기 위해, 본 논문에서는 합의 기반 통신 및 지식 공유(CCKS) 프레임워크를 제시합니다. CCKS는 에이전트가 합의를 통해 도출된 제약 조건에 따라 추천을 채택하고, 교사의 지침을 보다 효율적으로 따르도록 합니다. 이러한 메커니즘은 에이전트가 탐색과 경험 많은 교사로부터 학습하는 균형을 맞추어 전체적인 성능을 향상시킵니다. 핵심은 합의 모델 구축이며, 이를 위해 에이전트의 훈련 단계에서 로컬 관찰 데이터를 기반으로 콘트라스트 학습을 활용하여 합의 모델을 구축하는 방법을 제안합니다. 행동 선택 시, 에이전트는 합의 및 공유 지식을 바탕으로 각 행동에 점수를 매기고 선택합니다. CCKS는 플러그 앤 플레이 방식으로 설계되어 기존 DTDE 알고리즘과 원활하게 통합됩니다. Google Research Football 환경 및 복잡한 StarCraft II Multi-Agent Challenge에서 수행된 실험 결과, CCKS를 적용하면 협력 효율성, 학습 속도 및 전체 성능이 현재의 DTDE 기준선보다 크게 향상되는 것을 확인했습니다. 코드 내용은 다음 링크에서 확인할 수 있습니다: https://github.com/yuanxpy/CCKS.

Original Abstract

In Decentralized Training and Decentralized Execution (DTDE) for cooperative Multi-Agent Reinforcement Learning (MARL), action-advising-based knowledge sharing promotes interpretable and scalable cooperation among agents. However, current action advising approaches often adhere too much to the teacher's guidance without evaluating teacher-student compatibility, which causes excessive advising, suboptimal stability, and degraded performance. To overcome these challenges, this paper presents a Consensus-based Communication and Knowledge Sharing (CCKS) framework, which allows agents to adopt recommendations based on consensus-derived constraints and to follow the teacher's instructions more smartly. This mechanism enables agents to balance exploration and learning from experienced teachers, improving overall performance. The key is the consensus model construction, for which we propose to employ contrastive learning to construct consensus models based on local observations in the agents' training phase. In action selection, agents score and choose actions based on consensus and shared knowledge. Designed as a plug-and-play solution, CCKS integrates seamlessly with existing DTDE algorithms. Experiments conducted in the Google Research Football environment and the complex StarCraft II Multi-Agent Challenge demonstrate that the integration with CCKS significantly improves cooperation efficiency, learning speed, and overall performance compared with current DTDE baselines. The code is available at https://github.com/yuanxpy/CCKS.

0 Citations
0 Influential
27.5 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!