HAFI-VLM: 시각 언어 모델의 시각적 인지 진단 및 향상을 위한 주파수 기반 접근 방식
HAFI-VLM: A Frequency Perspective for Diagnosing and Enhancing Visual Perception in Vision-Language Models
시각 언어 모델(VLMs)은 예측에 미세한 시각 정보가 필요할 때 여전히 신뢰성이 떨어지는 경향이 있습니다. 본 연구에서는 이 문제의 간과되었던 원인, 즉 '스펙트럼 응답 강직성'을 밝혀냈습니다. 사전 학습된 시각 인코더는 이미지 및 작업 전반에 걸쳐 상당한 주파수 변화가 존재함에도 불구하고, 다운스트림 미세 조정 과정에서도 크게 변하지 않는 인코더별 레이어 단위의 일관된 스펙트럴 프로필을 나타냅니다. 사전 학습된 시각 인코더는 이미지만 입력받기 때문에 현재 쿼리에 필요한 증거에 맞춰 스펙트럼 추출을 적응시킬 수 없습니다. 따라서 본 연구에서는 사전 학습된 의미 표현을 유지하면서 작업 조건에 맞는 주파수 경로를 도입하는 HAFI-VLM을 제안합니다. 계층적 적응형 주파수 주입(HAFI)은 텍스트 변조 및 공간적으로 정렬된 크로스 어텐션을 사용하여 여러 인코더 깊이에서 상호 보완적인 저주파, 중주파, 고주파 정보를 검색합니다. 시각 풍부화 레이어 어댑터는 또한 표면 LLM 어텐션을 재보정하여 강화된 시각 토큰을 효과적으로 활용할 수 있도록 합니다. LLaVA-1.5 및 Qwen2.5-VL에 대한 실험 결과, 일반적인 VQA(Visual Question Answering), 텍스트 기반 이해 능력, 그리고 환각 현상에 대한 강건성 측면에서 기존의 표현 수준 향상 방법 및 대부분의 해상도 또는 크롭 기반 접근 방식보다 뛰어난 성능을 보였습니다. 메커니즘 분석 결과, HAFI는 작업 의존적인 스펙트럴 할당을 복원하면서 의미 있는 어텐션을 유지하며, 주파수 강화가 VLM 인지 능력을 향상시키는 효과적인 방법임을 입증했습니다.
Vision-language models (VLMs) remain unreliable when predictions require fine-grained visual evidence. We identify a previously overlooked cause: spectral response rigidity. Despite substantial frequency variation across images and tasks, pretrained vision encoders exhibit persistent, encoder-specific layerwise spectral profiles that change only marginally under downstream fine-tuning. Since pretrained vision encoders only receive images, they cannot adapt spectral extraction to the evidence required by the current query. We therefore propose HAFI-VLM, which introduces a task-conditioned frequency pathway while preserving the pretrained semantic representation. Hierarchical Adaptive Frequency Injection (HAFI) retrieves complementary low-, mid-, and high-frequency evidence at multiple encoder depths using text-modulated, spatially aligned cross-attention. A Visual Enrichment Layer Adapter further recalibrates shallow LLM attention to effectively utilize the enriched visual tokens. Experiments on LLaVA-1.5 and Qwen2.5-VL demonstrate consistent improvements in general VQA, text-rich understanding, and hallucination robustness, outperforming representation-level enhancement methods and most resolution- or cropping-based approaches without additional high-resolution encoding. Mechanistic analyses show that HAFI restores task-dependent spectral allocation while retaining semantic attention, establishing frequency enrichment as a distinct and effective route for improving VLM perception.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.