2511.19969v2 Nov 25, 2025 cs.AI

M$^3$Prune: 효율적인 다중 모드 다중 에이전트 검색 기반 생성 모델을 위한 계층적 통신 그래프 가지치기

M$^3$Prune: Hierarchical Communication Graph Pruning for Efficient Multi-Modal Multi-Agent Retrieval-Augmented Generation

Chen Chen
Chen Chen
Citations: 101
h-index: 6
Zijie Zhou
Zijie Zhou
Citations: 61
h-index: 5
Weizi Shao
Weizi Shao
Citations: 0
h-index: 0
Taolin Zhang
Taolin Zhang
Citations: 362
h-index: 10
Chengyu Wang
Chengyu Wang
School of Software Engineering, East China Normal University, Shanghai, China
Citations: 712
h-index: 14
Xiaofeng He
Xiaofeng He
Citations: 125
h-index: 7

최근 다중 모드 검색 기반 생성(mRAG) 기술의 발전은 외부 지식을 활용하여 다중 모드 대규모 언어 모델(MLLM)을 향상시키며, 효과적인 통신을 통해 여러 에이전트들의 집단 지성이 단일 모델보다 훨씬 뛰어난 성능을 발휘할 수 있음을 보여주었습니다. 하지만 기존 다중 에이전트 시스템은 상당한 토큰 오버헤드와 증가된 계산 비용을 내포하고 있어, 대규모 배포에 어려움을 초래합니다. 이러한 문제점을 해결하기 위해, 저희는 새로운 다중 모드 다중 에이전트 계층적 통신 그래프 가지치기 프레임워크인 M$^3$Prune을 제안합니다. 저희의 프레임워크는 서로 다른 모달리티 간의 중복된 연결을 제거하여 작업 성능과 토큰 오버헤드 사이의 최적의 균형을 달성합니다. 구체적으로, M$^3$Prune은 먼저 텍스트 및 시각 모달리티 내에서 그래프 희소화를 적용하여 작업을 해결하는 데 가장 중요한 연결을 식별합니다. 그 후, 이러한 핵심 연결을 사용하여 동적인 통신 토폴로지를 구성하고, 모달 간 그래프 희소화를 수행합니다. 마지막으로, 중복된 연결을 점진적으로 제거하여 더욱 효율적이고 계층적인 구조를 얻습니다. 일반적인 벤치마크 및 특정 분야의 mRAG 실험 결과는 저희 방법이 단일 에이전트 시스템과 강력한 다중 에이전트 mRAG 시스템 모두보다 우수한 성능을 보이며, 토큰 사용량을 크게 줄이는 것을 입증했습니다.

Original Abstract

Recent advancements in multi-modal retrieval-augmented generation (mRAG), which enhance multi-modal large language models (MLLMs) with external knowledge, have demonstrated that the collective intelligence of multiple agents can significantly outperform a single model through effective communication. Despite impressive performance, existing multi-agent systems inherently incur substantial token overhead and increased computational costs, posing challenges for large-scale deployment. To address these issues, we propose a novel Multi-Modal Multi-agent hierarchical communication graph PRUNING framework, termed M$^3$Prune. Our framework eliminates redundant edges across different modalities, achieving an optimal balance between task performance and token overhead. Specifically, M$^3$Prune first applies intra-modal graph sparsification to textual and visual modalities, identifying the edges most critical for solving the task. Subsequently, we construct a dynamic communication topology using these key edges for inter-modal graph sparsification. Finally, we progressively prune redundant edges to obtain a more efficient and hierarchical topology. Extensive experiments on both general and domain-specific mRAG benchmarks demonstrate that our method consistently outperforms both single-agent and robust multi-agent mRAG systems while significantly reducing token consumption.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!