2608.04454v1 Aug 05, 2026 cs.CV

글로벌 라우팅 집계 기술을 넘어: MoE 비전-언어 모델을 위한 단계 인식 전문가 병합

Beyond Global Routing Aggregation: Phase-Aware Expert Merging for MoE Vision-Language Models

Wuyang Zhang
Wuyang Zhang
Citations: 44
h-index: 3
Xiangwen Xia
Xiangwen Xia
Citations: 0
h-index: 0
Hongyu Zhang
Hongyu Zhang
Citations: 28
h-index: 3
Cheng Yan
Cheng Yan
Citations: 122
h-index: 3

혼합 전문가(MoE) 기반의 비전-언어 모델(VLM)은 희소한 전문가 활성화를 통해 모델 용량을 늘리지만, 배포 과정에서 전체 전문가 풀을 저장해야 하는 부담이 있습니다. 훈련 없이 전문가를 병합하는 기술은 이러한 부담을 줄여주며, 많은 라우팅 기반 방법들은 모든 토큰에 대한 라우팅 통계를 집계하여 병합 가능성을 판단합니다. 그러나 MoE-VLM 추론은 이미지 컨텍스트 토큰(시각적 정보 전달), 질문 토큰(쿼리 지정), 답변 토큰(결과 생성)의 세 단계로 구성되며, 각 단계는 서로 다른 수의 토큰과 라우팅 분포를 가집니다. 특히 이미지 컨텍스트 토큰이 훨씬 많기 때문에, 글로벌 집계 방식은 이미지 컨텍스트 처리에 지나치게 집중하고, 단계에 따라 다른 역할을 수행하는 전문가들의 역할 분이를 흐려 모델 성능을 저하시킬 수 있습니다. 따라서 우리는 MoE-VLM의 전문가 병합 과정에서 단계에 따른 전문가의 역할을 보존해야 하며, 글로벌 집계된 라우팅 통계가 아닌 각 전문가가 어떤 단계를 주로 처리하는지에 따라 병합 가능성을 판단해야 한다고 주장합니다. 이러한 관점에서, 우리는 단계별 라우팅 통계를 정규화하여 각 전문가의 '라우팅 역할 프로필(RRP)'을 생성하고, 이를 기반으로 전문가를 병합하는 훈련이 필요 없는 방법인 RoleMerge를 제안합니다. RoleMerge는 전문가-단계 정보를 활용하여 호환 가능한 전문가 프로필과 해당 라우터 항목을 함께 병합하면서 답변 디코딩에 사용되는 전문가들의 구분을 유지합니다. 세 가지 모델과 다양한 벤치마크에서 수행한 실험 결과, RoleMerge는 다른 전문가 병합 방법에 비해 더 높은 성능을 유지하며, 특히 여섯 가지 작업의 평균 성능에서 최대 9.6% 향상을 보였습니다. 이러한 결과는 단계에 따른 전문가 역할이 MoE-VLM 전문가 병합을 위한 글로벌 라우팅 집계보다 효과적인 기준임을 입증합니다.

Original Abstract

Mixture-of-experts vision-language models (MoE-VLMs) increase model capacity with sparse expert activation, yet deployment requires storing the full expert pool. Training-free expert merging reduces this burden, and many routing-based methods aggregate routing statistics across all tokens to determine merge compatibility. However, MoE-VLM inference is phase-structured: image-context tokens carry visual content, question tokens specify the query, and answer tokens produce the output, with different counts and routing distributions. Because image-context tokens are far more numerous, global aggregation can overemphasize image-context processing and obscure phase-conditioned expert roles, making experts serving different phases appear interchangeable and degrading model performance. We therefore argue that MoE-VLM expert merging should preserve phase-conditioned expert roles, judging compatibility by how experts serve different phases rather than globally aggregated routing statistics. Based on this view, we propose RoleMerge, a training-free method that constructs each expert's Routing Role Profile (RRP) from phase-normalized routing statistics, capturing its relative phase preference. Guided by expert-phase information loss, RoleMerge merges experts with compatible profiles and their corresponding router entries while preserving answer-decoding expert distinctions. Experiments on three models and multiple benchmarks show that RoleMerge preserves more of the full model's performance than alternative expert-merging methods at matched expert-retention ratios, with relative improvements of up to 9.6 percent in six-task macro-average performance. These results validate phase-conditioned expert roles as a more effective basis than global routing aggregation for MoE-VLM expert merging.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!