2607.29071v1 Jul 31, 2026 cs.LG

연합 학습 기반의 파운데이션 모델 미세 조정: 이기종 압축 클라이언트를 활용

Federated Foundation Models Fine-Tuning with Heterogeneous Compressed Clients

Mayi Xu
Mayi Xu
Citations: 45
h-index: 3
Shengkun Zhu
Shengkun Zhu
Citations: 56
h-index: 5
Jinshan Zeng
Jinshan Zeng
Citations: 12
h-index: 2
Zhihua Allen-Zhao
Zhihua Allen-Zhao
Citations: 31
h-index: 3
Quanqing Xu
Quanqing Xu
Citations: 47
h-index: 4
Wei Ren
Wei Ren
Citations: 0
h-index: 0
Qiang Yang
Qiang Yang
Citations: 250
h-index: 2
Yang Liu
Yang Liu
Citations: 247
h-index: 5

파운데이션 모델의 연합 학습은 근본적인 자원 비대칭 문제에 직면합니다. 즉, 가장 가치 있는 도메인 특화 데이터를 보유한 기관조차 수십억 개의 파라미터를 가진 모델을 호스팅할 수 없습니다. 기존의 이기종 연합 학습 접근 방식은 파라미터 효율적인 조정, 모델 가지치기 또는 지식 증류를 통해 이러한 격차를 해소하려고 시도하지만, 각 방법은 전체 모델 메모리 축소, 아키텍처 자체 완성성 또는 표현 충실성과 같은 중요한 특성을 희생하여 근본적인 긴장을 해결하지 못합니다. 우리는 파라미터 중심의 연합 미세 조정 프레임워크인 FedSLM을 제안합니다. FedSLM은 SVD 기반 분해를 사용하여 자체적으로 작동 가능한 클라이언트 모델을 생성하며, 이 모델들의 저차원 부분 공간은 구조적으로 호환되는 중첩된 다양체를 형성합니다. 그런 다음, FedSLM은 압축 그룹 내의 경량 어댑터를 동기화하고, 구조적 정렬을 통해 그룹 간에 전체 랭크 재구성을 결합하는 두 단계 프로토콜을 적용합니다. 마지막으로, 보조 신뢰도 손실을 사용한 약-강(weak-to-strong) 엘리시테이션 단계를 통해 집계된 지식을 전체 규모 서버로 전달하고, 명시적인 편향-분산 균형을 통해 압축으로 인한 왜곡을 완화합니다. 우리는 어댑터 수준의 집계에 대한 이론적 보장을 제공하고, 그룹 간 융합을 위한 부분 공간 정렬 경계를 제시하며, 신뢰도 손실이 약한 감독 데이터의 노이즈를 어떻게 완화하는지에 대한 분석을 제공합니다. 자연어 처리 및 이미지-텍스트 벤치마크 실험 결과, FedSLM은 IID 및 non-IID 분할 모두에서 기존 연합 학습 모델보다 우수한 성능을 보이며, 클라이언트 모델은 전체 모델에 필요한 GPU 메모리의 약 50%만 사용합니다.

Original Abstract

Federated learning of foundation models faces a fundamental resource-asymmetry challenge: the institutions holding the most valuable domain-specific data cannot host billion-parameter models. Existing heterogeneous federated approaches attempt to bridge this gap through parameter-efficient tuning, model pruning, or knowledge distillation, yet each trades away a critical property, whether full-model memory reduction, architectural self-containedness, or representational fidelity, leaving the core tension unresolved. We propose FedSLM, a parameter-centric framework for federated fine-tuning with heterogeneous compressed clients. FedSLM uses SVD-based decomposition to produce self-contained client models, whose low-rank subspaces form nested manifolds that are structurally compatible for aggregation. It then applies a two-stage protocol that synchronizes lightweight adapters within compression groups and fuses full-rank reconstructions across groups via structural alignment. Finally, a weak-to-strong elicitation step with auxiliary confidence loss transfers the aggregated knowledge to the full-scale server, while an explicit bias--variance trade-off mitigates compression artifacts. We provide theoretical guarantees for adapter-level aggregation, subspace-alignment bounds for cross-group fusion, and a characterization of how the confidence loss mitigates weak-supervision noise. Experiments on natural language and vision--language benchmarks show that FedSLM outperforms existing federated baselines under both IID and non-IID partitions, while client models operate at roughly 50% of the GPU memory required by the full model.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!