2605.04504v1 May 06, 2026 cs.CV

SpecPL: 프롬프트 학습을 위한 스펙트럼 세분화 분리

SpecPL: Disentangling Spectral Granularity for Prompt Learning

Jingtao Zhou
Jingtao Zhou
Citations: 15
h-index: 1
Xirui Kang
Xirui Kang
Citations: 142
h-index: 1
Feiyang Huang
Feiyang Huang
Citations: 5
h-index: 1
L. Po
L. Po
Citations: 7,511
h-index: 37

기존의 대규모 시각언어 모델(VLM) 프롬프트 학습 방법은 모달리티 비대칭성을 나타내며, 주로 텍스트 토큰을 최적화하는 반면, 전체적인 특징 추출기로 사용되는 시각 인코더는 여전히 고정되어 있으며, 미세한 차별화를 위해 필수적인 스펙트럼 세분화 수준을 간과합니다. 이를 해결하기 위해, 우리는 Counterfactual Granule Supervision을 통해 프롬프트 학습에 대한 새로운 스펙트럼 관점을 제시하는 Disentangling Spectral Granularity for Prompt Learning (SpecPL)을 소개합니다. 구체적으로, 우리는 고정된 VAE를 사용하여 시각 신호를 의미론적 저주파 대역과 세분화된 고주파 세부 정보로 분해합니다. 고정된 시각 의미론적 뱅크(Visual Semantic Bank)는 텍스트 표현을 보편적인 저주파 불변 특성에 연결하여 과적합을 완화합니다. 더욱 중요한 것은, 미세한 차별화는 Counterfactual Granule 훈련을 통해 이루어집니다. 고주파 신호를 재배열함으로써, 모델은 시각적 세분화와 의미론적 불변성을 명시적으로 구별하도록 강제합니다. SpecPL은 고유하게, 시각적인 가이드라인을 통해 CoOp 및 MaPLe와 같은 텍스트 기반 모델을 활성화하는 범용 플러그 앤 플레이 부스터 역할을 합니다. 11개의 벤치마크에 대한 실험 결과, 경쟁력 있는 최첨단 성능을 달성했으며, 81.51%의 조화 평균 정확도라는 새로운 성능 수준을 달성했습니다. 이러한 결과는 스펙트럼 분리와 Counterfactual Supervision이 안정성과 일반화 사이의 균형을 효과적으로 개선한다는 것을 입증합니다. 코드: https://github.com/Mlrac1e/SpecPL-Prompt-Learning

Original Abstract

Existing prompt learning for VLMs exhibits a modality asymmetry, predominantly optimizing text tokens while still relying on frozen visual encoder as holistic extractor and neglecting the spectral granularity essential for fine-grained discrimination. To bridge this, we introduce Disentangling Spectral Granularity for Prompt Learning (SpecPL), which approaches prompt learning from a novel spectral perspective via Counterfactual Granule Supervision. Specifically, we leverage a frozen VAE to decompose visual signals into semantic low-frequency bands and granular high-frequency details. A frozen Visual Semantic Bank anchors text representations to universal low-frequency invariants, mitigating overfitting. Crucially, fine-grained discrimination is driven by counterfactual granule training: by permuting high-frequency signals, we compel the model to explicitly distinguish visual granularity from semantic invariance. Uniquely, SpecPL serves as a universal plug-and-play booster, revitalizing text-oriented baselines like CoOp and MaPLe via visual-side guidance. Experiments on 11 benchmarks demonstrate competitive state-of-the-art performance, achieving a new performance ceiling of 81.51\% harmonic-mean accuracy. These results validate that spectral disentanglement with counterfactual supervision effectively bridges the gap in the stability-generalization trade-off. Code is released at https://github.com/Mlrac1e/SpecPL-Prompt-Learning.

0 Citations
0 Influential
51.324746787308 Altmetric
0.0 Score
Original PDF
12

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!