2605.28306v1 May 27, 2026 cs.CL

Mixture-of-Experts 모델에서 다국어 다운스트림 작업을 위한 라우팅 정렬 기반 미세 조정

Routing-Aligned Fine-Tuning for Multilingual Downstream Tasks in Mixture-of-Experts Models

Guanzhi Deng
Guanzhi Deng
Citations: 27
h-index: 3
Kuan Wu
Kuan Wu
Citations: 0
h-index: 0
Shing Yin Wong
Shing Yin Wong
Citations: 0
h-index: 0
Linqi Song
Linqi Song
Citations: 449
h-index: 7
Haibo Wang
Haibo Wang
Citations: 20
h-index: 4
Sichun Luo
Sichun Luo
Citations: 860
h-index: 12

Mixture-of-Experts (MoE) 모델은 효율적인 LLM 확장을 위한 주요 패러다임으로 부상했지만, 이를 비영어권의 다양한 작업에 적용하는 것은 여전히 어려운 과제입니다. 기존의 미세 조정 방법은 MoE 모델을 단일 학습자로 취급하여, 사전 훈련 중에 형성되는 이질적인 라우팅 구조를 고려하지 않습니다. 여러 MoE 모델과 다운스트림 작업을 통해 중간 레이어가 언어 보편적인 정렬 영역을 형성하며, 라우팅 발산 정도가 각 언어별 작업 성능 격차를 강력하게 예측한다는 것을 확인했습니다. 이러한 관찰을 바탕으로, 우리는 RA-MoE (Routing-Aligned MoE Fine-Tuning)라는 세 단계 프레임워크를 제안합니다. 이 프레임워크는 영어 및 대상 언어에서의 정확성을 기준으로 병렬 작업 예제를 4가지 유형(cc/ci/ic/ii)으로 분류하고, 중간 레이어에서 작업과 관련된 전문가를 식별하며, 표준 SFT에 라우팅 정렬 손실을 추가하여 대상 언어의 ci 유형 예제에 대한 라우팅이 영어 작업-전문가 활성화 패턴을 따르도록 장려합니다. 세 가지 MoE 모델, 세 가지 작업, 여섯 개의 대상 언어를 대상으로 한 실험 결과, RA-MoE는 표준 SFT 및 Routing Steering, RISE와 같은 강력한 기준 방법을 지속적으로 능가했으며, 작업-언어 쌍의 ci 비율이 정렬 효과를 예측하는 데 신뢰할 수 있는 지표임을 확인했습니다.

Original Abstract

Mixture-of-Experts (MoE) models have emerged as a dominant paradigm for efficient LLM scaling, yet adapting them to non-English downstream tasks remains challenging. Existing fine-tuning approaches treat MoE models as monolithic learners, ignoring the heterogeneous routing structure that develops during pretraining. We validate across multiple MoE models and downstream tasks that middle layers form a language-universal alignment zone where routing divergence strongly predicts per-language task performance gaps. Building on this observation, we propose RA-MoE (Routing-Aligned MoE Fine-Tuning), a three-stage framework that categorizes parallel task examples into a four-way taxonomy (cc/ci/ic/ii) based on correctness in English and the target language, identifies task-relevant experts in the middle layers, and augments standard SFT with a routing alignment loss that encourages target-language routing on ci-type examples to follow the English task-expert activation pattern. Experiments across three MoE models, three tasks, and six target languages demonstrate that RA-MoE consistently outperforms standard SFT and strong baselines including Routing Steering and RISE, with the ci proportion of a task-language pair serving as a reliable predictor of alignment benefit.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!