FedMPT: 분산 다중 레이블 프롬프트 튜닝을 통한 시각-언어 모델
FedMPT: Federated Multi-label Prompt Tuning of Vision-Language Models
시각-언어 모델(VLMs) 기반의 다중 레이블 인식(MLR)은 사전 학습된 지식을 활용하여 복잡한 인식 시나리오에 더 잘 적응하고, 모델의 견고성을 향상시키는 것을 목표로 합니다. 그러나 현실적인 분산 환경에서의 연합 학습(Federated Learning)을 위해서는 각 클라이언트가 보유한 개인적이고 이질적인 데이터를 사용하여 VLMs를 조정하는 과정에서 모델이 인위적인 레이블 상관관계를 과도하게 학습하여, 새로운 샘플을 접했을 때 관련 없는 범주를 활성화시키는 문제가 발생할 수 있습니다. 이러한 문제를 해결하기 위해, 우리는 원인 모델을 기반으로 한 연합 학습 MLR 방식을 재검토하며, 중간 변수를 통해 오라클 레이블의 공존 관계를 증폭시키는 방식으로 MLR 모델링 과정을 분리하는 프론트-도어 조정 기법을 도입합니다. 저희의 분석에 따라, 특별히 분산 MLR을 위해 설계된 첫 번째 방법인 FedMPT를 제안합니다. FedMPT의 핵심 아이디어는 일반화 가능한 조건을 활용하여 연합 MLR이 오작동하는 레이블 활성화를 완화하도록 유도하는 것입니다. 이를 달성하기 위해, FedMPT는 대규모 언어 모델(LLM) 기반 파이프라인을 도입하여 레이블 의존성을 규정하는 근본적인 조건을 파악합니다. 또한, 조건 정보를 풍부하게 담은 프롬프트와 이미지 패치 간의 최적 수송을 통해 다양한 영역 레벨의 의미를 밝혀냅니다. 마지막으로, 특별히 설계된 게이트 메커니즘을 사용하여 다양한 조건으로부터 시너지 효과를 내는 예측을 생성합니다. 여러 벤치마크 데이터셋에 대한 실험 결과, 저희가 제안하는 방법은 경쟁력 있는 성능을 달성하며, 다양한 환경에서 최첨단(SOTA) 방법보다 우수한 결과를 보여줍니다.
Multi-Label Recognition (MLR) based on Vision-Language Models (VLMs) aims to leverage their pre-trained knowledge to better adapt complex recognition scenarios, thereby enhancing model robustness. However, for realistic decentralized applications requiring federated learning, adapting VLMs to each client that possesses private and heterogeneous data can cause the model to overfit spurious label correlations, consequently triggering irrelevant categories when encountering new samples. To tackle this problem, we reconsider the federated learning for MLR with a causal model, in which we adopt a front-door adjustment and decouple the MLR modeling process by intermediate variables that magnify the oracle label co-occurrence. Guided by our analysis, we propose our FedMPT, the first method specifically designed for federated MLR. The core idea of FedMPT is to leverage generalizable conditions to steer federated MLR to mitigate erroneous label activations. To achieve this, FedMPT introduces an Large Language Model (LLM)-driven pipeline to decipher the underlying conditions that govern label dependencies. Furthermore, we introduce an optimal transport between the condition-enriched prompts and the image patches to uncover multiple region-level semantics. Finally, we generate synergistic predictions from different conditions with a crafted gating mechanism. Experiments on multiple benchmark datasets show that our proposed approach achieves competitive results and outperforms SOTA methods under varied settings.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.