2606.25389v1 Jun 24, 2026 cs.AI

기술 분할 및 재사용을 통한 오프라인 다중 에이전트 지속적인 협력

Offline Multi-agent Continual Cooperation via Skill Partition and Reuse

Ruiqi Xue
Ruiqi Xue
Citations: 69
h-index: 2
Lei Yuan
Lei Yuan
Nanjing University
Citations: 857
h-index: 16
Yang Yu
Yang Yu
Citations: 297
h-index: 5
Yuchen Xiao
Yuchen Xiao
Citations: 20
h-index: 2
Tieyue Yin
Tieyue Yin
Citations: 0
h-index: 0

다중 에이전트 오프라인 데이터셋에서 기술을 추출하는 것은 작업 불변의 조정 기술을 다양한 작업에 공유함으로써 학습 효율성을 향상시킵니다. 작업이 순차적으로 발생하고 기술 공간이 기하급수적으로 증가하는 환경에서, 휴리스틱하게 설계되고 고정 크기의 기술 라이브러리에 의존하는 기존 접근 방식은 데이터 분포 변화 및 간섭 문제를 해결하는 데 어려움을 겪으며, 파국적인 망각과 가소성 손실을 초래합니다. 이러한 문제를 해결하고 에이전트가 개방형 환경에서 지속적으로 조정 기술을 발견하고 재사용할 수 있는 능력을 부여하기 위해, 우리는 기술 분할 및 재사용을 통한 지속적인 오프라인 다중 에이전트 기술 발견을 위한 체계적인 프레임워크인 COMAD를 제안합니다. 먼저, 자기 부호화기를 사용하여 혼합된 다중 에이전트 행동 데이터에서 기술을 발견하고 조정 지식을 재사용 가능한 조정 기술로 변환합니다. 그런 다음, 밀도 기반의 재사용성 추정기를 통해 식별된 재사용 가능한 기술을 명시적으로 활용하여 장점 함수를 안내하는 멀티 헤드 아키텍처를 갖는 기술 증강 정책 학습 목표를 구성합니다. 이론적 분석에 따르면, 제안된 방법은 지속적인 기술 발견 문제의 최적값을 근사합니다. 다양한 MARL 벤치마크에서 얻은 실험 결과는 COMAD가 간섭을 완화하기 위해 지속적으로 기술 라이브러리를 확장하며, 여러 기준 모델보다 우수한 순방향 및 역방향 전이 성능을 달성한다는 것을 보여줍니다.

Original Abstract

Extracting skills from multi-agent offline dataset improves learning efficiency via sharing task-invariant coordination skills among tasks. In settings where tasks occur sequentially and the space of skills grows exponentially, existing approaches that rely on heuristically designed and fixed-sized skill libraries struggle to resolve the problem of distributional shift and interference, facing catastrophic forgetting and plasticity loss. To address this problem and endow agents with the ability to continually discover and reuse coordination skills in open-environment, we propose COMAD, a principled framework for Continual Offline Multi-agent Skill Discovery via Skill Partition and Reuse. We first discover skills from mixed multi-agent behavior data with an auto-encoder to transform coordination knowledge into reusable coordination skills. Then we construct a skill-augmented policy learning objective with multi-head architectures, explicitly guiding the advantage function with reusable skills identified via a density-based reusability estimator. Theoretical analysis shows our method approximates the optimum of a continual skill discovery problem. Empirical results across diverse MARL benchmarks show that COMAD continually expands its skill library to mitigate interference, achieving superior forward and backward transfer for task streams compared to multiple baselines.

1 Citations
0 Influential
8 Altmetric
41.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!