2606.31413v1 Jun 30, 2026 cs.AI

재학습이 아닌 선택을 배우는 방법: 하드 라우팅 기반 추론 LoRA 혼합 모델

Learning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAs

Prayag Tiwari
Prayag Tiwari
Citations: 75
h-index: 4
Zhan Su
Zhan Su
Citations: 42
h-index: 4
Seyed Alireza Molavi
Seyed Alireza Molavi
Citations: 6
h-index: 1
Yan Hu
Yan Hu
Citations: 104
h-index: 5
Peyman Sheikholharam Mashhadi
Peyman Sheikholharam Mashhadi
Citations: 4
h-index: 1
Stefan Byttner
Stefan Byttner
Citations: 10
h-index: 2

독립적으로 학습된 LoRA 어댑터를 하나의 대규모 언어 모델로 결합하는 것은 특히 원래 학습 데이터를 공유할 수 없을 때 다중 도메인 적응에 유용합니다. 일반적인 방법은 MoE(Mixture of Experts) 스타일의 라우팅을 사용하여 LoRA 전문가를 조합하는 것입니다. 그러나 동결된 사전 훈련된 어댑터의 경우, 소프트 가중치 결합은 각 LoRA 모듈이 원래 학습되었던 단위 스케일 추가 업데이트 방식을 변경할 수 있습니다. 본 논문에서는 하드 라우팅 MoR-LoRA(Hard-Routed Mixture of Reasoning LoRAs)라는 두 단계 프레임워크를 제안합니다. 이 프레임워크는 검증 가능한 피드백을 통해 강화 학습을 사용하여 독립적으로 훈련된 도메인별 LoRA 어댑터를 추론 전문가로 확보하는 방식으로 작동합니다. 그런 다음 모든 전문가를 동결하고, 그들의 추론 과정을 추출하여, 경량화된 공유 라우터와 작은 Attention LoRA만을 결합을 위해 훈련합니다. 라우터는 하드 Top-1 라우팅을 사용하여 토큰당 정확히 하나의 전문가를 선택하며, Straight-Through Estimator는 기울기 기반 학습을 가능하게 합니다. 다섯 가지 벤치마크, 다양한 모델 크기 및 추가 모델 패밀리에 대한 실험 결과, Hard-Routed MoR-LoRA는 전문가의 동작을 유지하면서 소프트 라우팅 기반 모델보다 훨씬 적은 수의 학습 가능한 파라미터가 필요함을 보여줍니다. 또한 분석 결과, 정규화된 소프트 혼합은 종종 대부분의 라우팅 가중치를 단일 전문가에 집중시키는 경향이 있으며, 이는 하드 단위 스케일 라우팅이 동결된 LoRA 전문가 결합을 위한 간단하고 효율적인 추상화 방법임을 시사합니다.

Original Abstract

Composing independently trained LoRA adapters into a single large language model is useful for multi-domain adaptation, especially when the original training data cannot be shared. A common approach is to use MoE-style routing over LoRA experts, but for frozen pretrained adapters, soft weighted combinations can change the unit-scale additive update under which each LoRA module was originally trained. We propose \textbf{Hard-Routed MoR-LoRA}, a two-stage framework for composing frozen reasoning LoRA experts through unit-scale hard selection. First, domain-specific LoRA adapters are trained independently using reinforcement learning from verifiable feedback to obtain reasoning experts. Then, all experts are frozen, reasoning traces are distilled from them, and only a lightweight shared router together with a small attention LoRA is trained for integration. The router selects exactly one expert per token using hard top-1 routing, while a straight-through estimator enables gradient-based training. Experiments across five benchmarks, multiple model scales, and additional model families show that Hard-Routed MoR-LoRA preserves expert behavior while requiring substantially fewer trainable parameters than soft-routing mixture baselines. Our analysis further shows that normalized soft mixtures often concentrate most routing mass on a single expert, suggesting that hard unit-scale routing provides a simple and efficient abstraction for frozen LoRA expert composition.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!