2606.23607v1 Jun 22, 2026 cs.LG

선형 모드 연결 및 병합을 활용한 대규모 사전 학습 트랜스포머 모델 확장

Scaling Linear Mode Connectivity and Merging to Billion Parameter Pretrained Transformers

Tianyi Li
Tianyi Li
Citations: 793
h-index: 5
Zhiqiang Shen
Zhiqiang Shen
Citations: 108
h-index: 3

선형 모드 연결(LMC)은 독립적으로 학습된 신경망을 이해하고 결합하는 데 유망한 기반을 제공하지만, 기존 방법들은 일반적으로 하나의 모델 지점에서의 보간 경로를 최적화하기 때문에 대규모 사전 학습 트랜스포머에 대한 확장성과 효과성이 제한됩니다. 본 연구에서는 { }대규모 사전 학습 트랜스포머 모델의 LMC 기반 병합을 가능하게 하는 새롭고 확장 가능한 프레임워크를 제안합니다. 저희 방법은 기능적으로 동등한 솔루션을 정렬하기 위해 적절하게 매개변수화된 기능 보존 가중치 변환을 적용하고, 두 모델이 공유되는 선형 보간 경로를 향해 해당 변환을 공동으로 학습하는 이중 학습 절차를 도입합니다. 이러한 양방향 최적화는 보간 장벽을 크게 줄이고 대규모 아키텍처에서의 더 안정적인 병합을 가능하게 합니다. 실험적으로, 저희 방법은 중간 크기의 매개변수를 가진 언어 모델에서 WikiText 데이터셋에 대해 거의 제로 수준의 손실 장벽을 달성했으며, 이는 현재까지 이 규모에서 거의 장벽 없는 선형 연결성을 보여주는 최초의 사례입니다. 컴퓨터 비전 분야에서는 ViT-L 모델이 보간 경로 전체에서 69% 이상의 ImageNet 상위 1개 정확도를 유지하는 반면, 최신 대규모 언어 모델은 작은 손실 장벽만을 나타냅니다. 이러한 결과는 적절하게 매개변수 대칭성을 해결하면 대규모 사전 학습 트랜스포머가 간단한 선형 경로를 통해 연결되고 병합될 수 있으며, 이를 통해 보간 성능이 크게 향상된다는 것을 시사합니다. 코드: https://github.com/VILA-Lab/Dual-Learned-Matching

Original Abstract

Linear mode connectivity (LMC) provides a promising foundation for understanding and merging independently trained neural networks, but existing methods typically optimize the interpolation path from only one model endpoint, limiting their scalability and effectiveness for large pretrained transformers. We propose a novel and scalable framework for enabling LMC-based model merging to {\em billion-parameter pretrained transformers}. Our method applies properly parameterized functionality-preserving weight transformations to align functionally equivalent solutions, and introduces a dual learning procedure in which both models jointly learn their corresponding transformations toward a shared linear interpolation path. This bidirectional optimization substantially reduces interpolation barriers and enables more reliable merging across large-scale architectures. Empirically, we show that our approach achieves near-zero loss barriers on WikiText for language models with medium-sized parameters, representing, to our knowledge, the first demonstration of near-barrier-free linear connectivity at this scale. In the vision domain, ViT-L maintains above 69\% ImageNet top-1 accuracy throughout the interpolation path, while modern billion-parameter LLMs exhibit only small loss barriers. These results suggest that properly resolving parameter symmetries enables large pretrained Transformers to be connected and merged through simple linear paths with substantially improved interpolation performance. Code: https://github.com/VILA-Lab/Dual-Learned-Matching .

0 Citations
0 Influential
25.9657359028 Altmetric
0.0 Score
Original PDF
1

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!