스펙트럼 분리 및 향상: 표현 학습을 위한 이중 도메인 대비 프레임워크
Spectral Disentanglement and Enhancement: A Dual-domain Contrastive Framework for Representation Learning
최근 대규모 다중 모드 대비 학습은 풍부하고 전이 가능한 표현을 학습하는 데 상당한 성공을 거두었지만, 여전히 특징 차원의 균일한 처리와 학습된 특징의 고유한 스펙트럴 구조에 대한 간과라는 근본적인 한계를 가지고 있습니다. 경험적 증거에 따르면 고차원 임베딩은 종종 좁은 원뿔 형태로 붕괴되어, 중요한 작업 관련 의미를 작은 부분 공간에 집중시키는 반면, 대부분의 차원은 노이즈와 거짓 상관 관계로 채워집니다. 이러한 스펙트럴 불균형과 얽힘은 모델의 일반화 능력을 저해합니다. 본 논문에서는 임베딩 공간의 기하학적 구조와 그 스펙트럴 특성 간의 간극을 해소하는 새로운 프레임워크인 Spectral Disentanglement and Enhancement (SDE)를 제안합니다. 우리의 접근 방식은 특이값 분해를 활용하여 특징 차원을 작업에 중요한 의미를 포착하는 강력한 신호, 부가적인 상관관계를 반영하는 약한 신호, 그리고 관련 없는 교란을 나타내는 노이즈로 적응적으로 분할합니다. 이후, 이론적인 안정성 보장을 제공하는 커리큘럼 기반의 스펙트럴 향상 전략을 적용하여 유용한 구성 요소를 선택적으로 증폭합니다. 향상된 특징을 기반으로, 우리는 특징 공간과 스펙트럴 공간 모두에서 정렬을 동시에 최적화하는 이중 도메인 대비 손실 함수를 추가로 도입하여, 스펙트럴 정규화를 학습 과정에 효과적으로 통합하고 더 풍부하고 강력한 표현을 장려합니다. 대규모 다중 모드 벤치마크에 대한 광범위한 실험 결과, SDE는 표현의 강건성과 일반화 능력을 지속적으로 향상시키며, 최첨단 방법보다 우수한 성능을 보임을 보여줍니다. SDE는 기존의 대비 학습 파이프라인과 원활하게 통합되어 다중 모드 표현 학습을 위한 효과적인 솔루션을 제공합니다.
Large-scale multimodal contrastive learning has recently achieved impressive success in learning rich and transferable representations, yet it remains fundamentally limited by the uniform treatment of feature dimensions and the neglect of the intrinsic spectral structure of the learned features. Empirical evidence indicates that high-dimensional embeddings tend to collapse into narrow cones, concentrating task-relevant semantics in a small subspace, while the majority of dimensions remain occupied by noise and spurious correlations. Such spectral imbalance and entanglement undermine model generalization. We propose Spectral Disentanglement and Enhancement (SDE), a novel framework that bridges the gap between the geometry of the embedded spaces and their spectral properties. Our approach leverages singular value decomposition to adaptively partition feature dimensions into strong signals that capture task-critical semantics, weak signals that reflect ancillary correlations, and noise representing irrelevant perturbations. A curriculum-based spectral enhancement strategy is then applied, selectively amplifying informative components with theoretical guarantees on training stability. Building upon the enhanced features, we further introduce a dual-domain contrastive loss that jointly optimizes alignment in both the feature and spectral spaces, effectively integrating spectral regularization into the training process and encouraging richer, more robust representations. Extensive experiments on large-scale multimodal benchmarks demonstrate that SDE consistently improves representation robustness and generalization, outperforming state-of-the-art methods. SDE integrates seamlessly with existing contrastive pipelines, offering an effective solution for multimodal representation learning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.