이질적인 대규모 언어 모델 병합에 대한 재고: 가중 모델 평균화 관점
Rethinking Heterogeneous LLM Merging: A Weighted Model Averaging Perspective
상당히 다른 파라미터 공간을 가진 대규모 언어 모델을 학습이나 의미적 정렬 없이 직접적인 가중 평균화를 통해 병합할 수 있을까요? 기존의 이질적인 융합 방법은 일반적으로 지식 증류, 어댑터, 학습된 잠재 공간, 라우팅 또는 특징 정렬과 같은 기술을 사용합니다. 본 연구에서는 진정으로 다른 파라미터를 가진 대규모 모델에 대해 더 간단한 방법이 효과적인지 탐구합니다. 우리는 학습 과정 없이 차원 조정을 수행하고 비율을 제어하여 모델을 보간하는 방식을 통해 이 직관에 반하는 질문을 재검토했습니다. '합집합' 방식의 병합에서는 작은 모델을 더 큰 파라미터 공간으로 확장하고, '교집합' 방식의 병합에서는 더 큰 모델을 더 작은 파라미터 공간으로 축소합니다. 수학적 추론, 코드 생성, 언어 이해, 상식 추론, 지식 및 명령어 따르기를 포함하는 다양한 벤치마크와 Qwen 패밀리의 모델 쌍을 사용하여 실험한 결과, 결정적인 확장은 원본 모델의 기능을 대부분 유지하며, 작은 비율로 보간하면 기존의 강력한 모델보다 더 나은 성능을 보여줄 수 있습니다. 그러나 거의 균형 잡힌 보간은 종종 실패하며, 작업 수준의 결과를 통해 특정 능력에서는 향상이 있지만 다른 능력에서는 성능이 저하되는 'seesaw 효과'가 나타납니다. 이러한 결과는 간단한 파라미터 평균화가 경량 차원 조정 및 신중하게 제어된 비율과 함께 이질적인 대규모 언어 모델 병합을 위한 놀랍도록 강력한 기준이 될 수 있음을 보여줍니다. 이는 직접적인 가중 융합의 한계가 더 복잡한 이질적인 융합 방법들이 달성할 수 있는 수준에도 영향을 미칠 수 있다는 점을 시사합니다.
Can large language models with substantially different parameter spaces be merged by direct weighted averaging, without training or semantic alignment? Existing heterogeneous fusion methods typically introduce distillation, adapters, learned latent spaces, routing, or feature alignment, leaving open whether a simpler recipe can work for genuinely different billion-parameter checkpoints. We revisit this counterintuitive question through training-free dimensional adaptation followed by ratio-controlled interpolation. In union-style merging, we expand the smaller model into the larger parameter space; in intersection-style merging, we truncate the larger model into the smaller parameter space. Across Qwen-family model pairs and benchmarks covering mathematical reasoning, code generation, language understanding, commonsense reasoning, knowledge, and instruction following, deterministic expansion largely preserves the source model function, and small-ratio interpolation can improve over strong source checkpoints by transferring complementary capabilities. However, near-balanced interpolation often collapses, and task-level results reveal a seesaw effect in which gains on some capabilities coexist with regressions on others. These results show that simple parameter averaging, when paired with lightweight dimensional adaptation and carefully controlled ratios, is a surprisingly strong baseline for heterogeneous LLM merging, suggesting that the limits of direct weighted fusion may also bound what more complex heterogeneous merging methods can achieve at scale.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.