이질적이고 불완전한 다중 모드 클라이언트 데이터를 활용한 연합 프롬프트 튜닝
Federated Prompt-Tuning with Heterogeneous and Incomplete Multimodal Client Data
본 논문에서는 현실적인 시나리오에서 로컬 데이터셋이 다중 모드이며 입력 수준에서 결측 특성의 분포 패턴이 서로 다른 경우를 고려한 일반화된 연합 프롬프트 튜닝 프레임워크를 제안합니다. 제안된 프레임워크는 기존의 단일 모드 또는 중앙 집중식 데이터에 초점을 맞추던 연합 학습과 다중 모드 프롬프트 튜닝 간의 격차를 해소합니다. 본 설정에서 핵심적인 과제는 서로 다른 클라이언트 간에 유사한 결측 데이터 분포 패턴을 인코딩하는 프롬프트 지침 간의 의미적 정렬 부족입니다. 이를 해결하기 위해, 본 논문에서는 각 클라이언트의 특성을 고려한 튜닝과 서버 측 집계 설계를 도입하여 프롬프트 튜닝 지침을 클라이언트와 데이터 모드 전반에 걸쳐 동시에 최적화하고 정렬하며 집계합니다. 이를 통해 프롬프트 지침이 서로 보완하고 효과적으로 결합될 수 있습니다. 다양한 다중 모드 벤치마크 데이터셋에 대한 광범위한 실험 결과는 본 연구가 최첨단(SOTA) 기준 모델보다 일관되게 우수한 성능을 보임을 입증합니다.
This paper introduces a generalized federated prompt-tuning framework for practical scenarios where local datasets are multi-modal and exhibit different distributional patterns of missing features at the input level. The proposed framework bridges the gap between federated learning and multi-modal prompt-tuning which have traditionally focused on either uni-modal or centralized data. A key challenge in this setting arises from the lack of semantic alignment between prompt instructions that encode similar distributional patterns of missing data across different clients. To address this, our framework introduces specialized client-tuning and server-aggregation designs that simultaneously optimize, align, and aggregate prompt-tuning instructions across clients and data modalities. This allows prompt instructions to complement one another and be combined effectively. Extensive evaluations on diverse multimodal benchmark datasets demonstrate that our work consistently outperforms state-of-the-art (SOTA) baselines.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.