2602.13309v1 Feb 10, 2026 cs.MA

적응적 가치 분해: 도시 시스템에서 변화하는 수의 에이전트 조정

Adaptive Value Decomposition: Coordinating a Varying Number of Agents in Urban Systems

Yexin Li
Yexin Li
Citations: 0
h-index: 0
Jinjin Guo
Jinjin Guo
Citations: 14
h-index: 2
Haoyu Zhang
Haoyu Zhang
Citations: 14
h-index: 2
Yuhan Zhao
Yuhan Zhao
Citations: 373
h-index: 9
Yiwen Sun
Yiwen Sun
Citations: 61
h-index: 4
Zihao Jiao
Zihao Jiao
Citations: 14
h-index: 1

다중 에이전트 강화 학습(MARL)은 다중 에이전트 시스템(MAS)을 조정하는 유망한 패러다임을 제공합니다. 그러나 대부분의 기존 방법은 고정된 에이전트 수 및 완전 동기화된 동작 실행과 같은 제한적인 가정에 의존합니다. 이러한 가정은 시간이 지남에 따라 활성 에이전트 수가 변하고, 동작의 지속 시간이 이질적인 도시 시스템에서 종종 위반되며, 이는 반-MARL 환경을 야기합니다. 또한, 에이전트 간 정책 매개변수를 공유하여 학습 효율성을 향상시키는 것이 일반적이지만, 유사한 관찰 환경에서 일부 에이전트가 동시에 의사 결정을 내릴 때 매우 균일한 동작을 초래하여 조정 품질을 저하시킬 수 있습니다. 이러한 과제를 해결하기 위해, 우리는 동적으로 변화하는 에이전트 집단에 적응하는 협력 MARL 프레임워크인 적응적 가치 분해(AVD)를 제안합니다. AVD는 또한 공유된 정책으로 인해 발생하는 동작 균질화를 완화하는 경량 메커니즘을 통합하여 에이전트 간 행동 다양성을 촉진하고 효과적인 협력을 유지합니다. 또한, 일부 에이전트가 다른 시간에 동작하는 비동기적 의사 결정을 수용할 수 있는 반-MARL 환경에 적합한 훈련-실행 전략을 설계했습니다. 런던과 워싱턴 D.C.의 두 주요 도시에서 실제 자전거 공유 재분배 작업에 대한 실험 결과, AVD는 최첨단 기준 모델보다 우수한 성능을 보이며, 그 효과성과 일반화 가능성을 입증합니다.

Original Abstract

Multi-agent reinforcement learning (MARL) provides a promising paradigm for coordinating multi-agent systems (MAS). However, most existing methods rely on restrictive assumptions, such as a fixed number of agents and fully synchronous action execution. These assumptions are often violated in urban systems, where the number of active agents varies over time, and actions may have heterogeneous durations, resulting in a semi-MARL setting. Moreover, while sharing policy parameters among agents is commonly adopted to improve learning efficiency, it can lead to highly homogeneous actions when a subset of agents make decisions concurrently under similar observations, potentially degrading coordination quality. To address these challenges, we propose Adaptive Value Decomposition (AVD), a cooperative MARL framework that adapts to a dynamically changing agent population. AVD further incorporates a lightweight mechanism to mitigate action homogenization induced by shared policies, thereby encouraging behavioral diversity and maintaining effective cooperation among agents. In addition, we design a training-execution strategy tailored to the semi-MARL setting that accommodates asynchronous decision-making when some agents act at different times. Experiments on real-world bike-sharing redistribution tasks in two major cities, London and Washington, D.C., demonstrate that AVD outperforms state-of-the-art baselines, confirming its effectiveness and generalizability.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!