SARA: 의미적으로 고정된 라우팅 정렬을 통한 전문가 모델의 다국어 지식 활용
SARA: Unlocking Multilingual Knowledge in Mixture-of-Experts via Semantically Anchored Routing Alignment
희소 전문가 모델(MoE) 아키텍처는 파라미터 확장성과 계산 효율성 사이의 균형을 제공하여 점점 더 중요한 패러다임으로 부상하고 있습니다. 그러나 고품질 훈련 데이터 부족으로 어려움을 겪는 저자원 언어는 종종 고자원 입력에 의해 주로 활성화되는 전문가와 다른 전문가에게 토큰이 라우팅되어, 다국어 환경에서의 효능을 제한합니다. 이러한 다국어 라우팅의 차이를 해결하기 위해, 우리는 SARA(Semantically Anchored Routing Alignment)라는 프레임워크를 제안합니다. SARA는 고자원 언어에서 얻은 전문적인 능력을 저자원 언어로 전달하는 앵커 역할을 수행하도록 설계되었습니다. SARA는 대칭 Jensen-Shannon (JS) 발산 제약을 사용하여 다국어 입력의 라우팅 분포를 고자원 의미적 앵커와 명시적으로 정렬합니다. 기존의 출력 로짓을 사용하는 증류 방법과 달리, SARA는 MoE 레이어의 내부 라우팅 분포를 직접 정렬하여 언어 간 전문가 선택의 메커니즘 일관성을 장려합니다. 우리는 2개의 LLM 모델을 사용하여 5개의 저자원 언어와 3가지 벤치마크에서 실험을 수행했습니다. 실험 결과는 SARA가 표준 인스트럭션 튜닝보다 우수한 성능을 보임을 보여줍니다 (예: Qwen3-30B-A3B에서 +0.8%, Phi-3.5-MoE-instruct에서 +1.2% 향상, Global-MMLU 기준). 추가 분석 결과는 SARA가 저자원 언어의 성능 병목 현상을 효과적으로 해결하며, 희소 아키텍처에서 다국어 기능을 향상시키는 확장 가능한 방법을 제공한다는 것을 보여줍니다.
Sparse Mixture-of-Experts (MoE) architectures have emerged as an increasingly influential paradigm as they offer a strategic balance between parameter scalability and computational efficiency. However, low-resource languages, which suffer from a scarcity of high-quality training data, often have their tokens routed to different experts than those predominantly activated by high-resource inputs, which limits cross-lingual expert sharing. This cross-lingual routing divergence consequently hinders their efficacy in multilingual contexts. To address this issue, we propose SARA (Semantically Anchored Routing Alignment), a framework designed to transfer specialized capabilities from high-resource languages as anchors to low-resource languages. SARA explicitly aligns the routing distribution of multilingual inputs with high-resource semantic anchors using a symmetric Jensen-Shannon (JS) divergence constraint. Unlike traditional distillation methods that operate on output logits, SARA directly aligns the internal routing distributions of MoE layers, encouraging mechanistic consistency in expert selection across languages. We conduct experiments on 2 LLMs across 5 low-resource languages and 3 benchmarks. Experiment results demonstrate that SARA outperforms standard instruction tuning, e.g., +0.8% on Qwen3-30B-A3B and +1.2% on Phi-3.5-MoE-instruct on Global-MMLU. Further analyses show that SARA effectively addresses performance bottlenecks in low-resource languages, providing a scalable pathway to enhance multilingual capabilities in sparse architectures.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.