L2R: 저랭크 및 립시츠 제어 라우팅을 이용한 믹스처 오브 Эксперт스
L2R: Low-Rank and Lipschitz-Controlled Routing for Mixture-of-Experts
믹스처 오브 Эксперт스 (MoE) 모델은 조건부로 활성화되는 소수의 Эксперт스를 활용하여 신경망의 규모를 확장하며, 이 과정에서 라우터는 Эксперт의 전문화 및 전체 모델 성능을 결정하는 핵심적인 역할을 수행합니다. 그러나 많은 최신 MoE 시스템은 여전히 고차원 표현 공간에서 선형 라우터를 사용하며, 이로 인해 표현 불일치, 각도 집중, 스케일에 민감한 스코링 등의 문제가 발생하여 라우팅의 분별력과 Эксперт의 안정적인 전문화를 저해할 수 있습니다. 본 연구에서는 라우팅 공간과 스코링 형상을 모두 재구성하는 통합 라우팅 프레임워크인 Low-rank & Lipschitz-controlled Routing (L2R)을 제안합니다. L2R은 공유된 저랭크 잠재 라우팅 공간에서 Эксперт 할당을 수행하며, 라우팅 함수의 립시츠 특성을 명시적으로 제어하는 Saturated Inner-Product Scoring (SIPS)를 도입하여 더 부드럽고 안정적인 라우팅 형상을 제공합니다. 또한, L2R은 Экспер트의 표현력을 향상시키는 파라미터 효율적인 멀티-앵커 라우팅 메커니즘을 포함하고 있습니다. 대규모 언어 MoE 모델 및 ImageNet 데이터셋을 사용한 비전 MoE 환경에서의 광범위한 실험 결과, L2R은 라우팅 안정성, Эксперт 전문화 및 전체 모델 성능을 지속적으로 향상시키는 것으로 나타났습니다.
Mixture-of-Experts (MoE) models scale neural networks by conditionally activating a small subset of experts, where the router plays a central role in determining expert specialization and overall model performance. However, many modern MoE systems still adopt linear routers in raw high-dimensional representation spaces, where representation mismatch, angular concentration, and scale-sensitive scoring can jointly undermine routing discriminability and stable expert specialization. In this work, we propose Low-rank \& Lipschitz-controlled Routing (L2R), a unified routing framework that reshapes both the routing space and scoring geometry. L2R performs expert assignment in a shared low-rank latent routing space and introduces Saturated Inner-Product Scoring (SIPS) to explicitly control the Lipschitz behavior of routing functions, yielding smoother and more stable routing geometry. In addition, L2R incorporates a parameter-efficient multi-anchor routing mechanism to enhance expert expressiveness. Extensive experiments on a large-scale language MoE model and a vision MoE setting on ImageNet demonstrate that L2R consistently improves routing stability, expert specialization, and overall model performance.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.