2605.28042v1 May 27, 2026 cs.CL

LLM에서 적극적인 전문가 제거를 통해 소규모 번역 전문 모델 추출

Extracting Small Translation Specialists from LLMs by Aggressively Pruning Experts

Lucas Bandarkar
Lucas Bandarkar
UCLA
Citations: 420
h-index: 6
Liu O. Martin
Liu O. Martin
Citations: 0
h-index: 0
Nanyun Peng
Nanyun Peng
Citations: 36
h-index: 4

최신 대형 언어 모델(LLM)은 뛰어난 기계 번역 성능을 달성하지만, 이는 주로 번역과 관련 없는 다양한 작업 및 기능에 대해 훈련된 광범위한 범용 모델이기 때문에 발생합니다. 따라서 이러한 모델은 번역이라는 특정 작업에는 과도하게 많은 파라미터를 가지고 있어, 상당한 메모리 및 계산 리소스가 필요합니다. 본 논문에서는 최신 Mixture-of-Experts (MoE) LLM에서 전문가를 적극적으로 제거하는 방법을 제시하며, 이 과정에서 번역 품질의 저하를 최소화합니다. 저희의 접근 방식은 전문가의 특수성과 LLM 내의 다국어 능력 분리성을 활용하여 번역에 관련 없는 전문가를 식별합니다. 또한 MoE 모델의 모듈식 구조 덕분에 이러한 전문가들은 별도의 훈련 없이 쉽게 제거될 수 있습니다. 재훈련 없이도, 저희는 전체 전문가의 절반을 제거하면서 미미한 품질 저하만 발생시키고, 70%까지 제거하며 경미한 손실만을 초래합니다. 짧은 Supervised Fine-Tuning (SFT) 과정을 통해 75%의 전문가를 제거하고 기본 성능을 복구할 수 있으며, 일부 환경에서는 거의 90%까지 제거하면서도 합리적인 번역 품질을 유지할 수 있습니다. 전반적으로 저희의 결과는 번역에 필요한 LLM의 파라미터가 실제 필요량보다 훨씬 많다는 것을 보여주며, 이는 90% 이상의 파라미터를 포함하는 MoE 블록의 상당한 압축을 가능하게 합니다.

Original Abstract

Modern large language models (LLMs) achieve state-of-the-art machine translation performance, but they do so as broad generalists largely trained for many tasks and capabilities unrelated to translation. Thus, they are heavily overparameterized for this task, resulting in excessive memory and compute requirements. In this paper, we present a method for aggressively pruning experts from modern mixture-of-experts LLMs while incurring negligible degradation in translation quality. Our approach exploits expert specialization and the separability of multilingual capabilities in LLMs to identify experts irrelevant to translation. And because of the modular nature of MoEs, these can be easily pruned without any training. Without retraining, we are able to prune half of all experts with negligible degradation and 70% with only minor losses. With a very short SFT, we prune 75% of experts while recovering baseline performance, and in some settings remove nearly 90% while maintaining reasonable translation quality. Overall, our results show that translation requires only a fraction of the LLM, enabling substantial compression of the MoE blocks that contain over 90% of parameters.

2 Citations
1 Influential
3 Altmetric
19.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!