2607.26107v1 Jul 28, 2026 cs.CV

TraceCLIP: 패치-CLS 기여도를 활용한 로컬 의미 정보 복원

TraceCLIP: Recovering Local Semantics from Patch-to-CLS Contributions

Xinran Liu
Xinran Liu
Citations: 202
h-index: 8
Shouqian Shi
Shouqian Shi
Citations: 635
h-index: 12
Ge Wang
Ge Wang
Citations: 0
h-index: 0
Xin-Wei Yao
Xin-Wei Yao
Citations: 0
h-index: 0
Yutong Chen
Yutong Chen
Citations: 0
h-index: 0
Sheng Zhong
Sheng Zhong
Citations: 0
h-index: 0

객체 위치 파악, 영역 인식 및 개방형 어휘 기반 의미 분할 등 밀집적인 시각-언어 이해는 언어 개념을 공간적으로 연결된 시각 영역과 연관시키는 것을 필요로 합니다. CLIP은 대규모 대비 학습을 통해 공유되는 이미지-텍스트 임베딩 공간을 학습함으로써 이러한 작업에 강력한 기반을 제공합니다. 그러나 CLIP의 이미지 레벨 목표는 텍스트를 CLS에서 파생된 전역 표현과 연결시키므로, 로컬 시각-언어 대응 관계는 간접적으로만 제약됩니다. 기존 방법들은 추가적인 지도, 외부 모델 또는 특정 작업에 맞는 적응을 도입하거나, 학습이 필요 없는 접근 방식은 주로 기존 패치 특징으로부터 밀집적인 응답을 복원하지만, CLIP 내에서 로컬 의미 정보가 가장 잘 드러나는 위치를 탐색하지 않습니다. 본 논문에서는 학습 없이 작동하는 TraceCLIP 프레임워크를 제안합니다. TraceCLIP은 CLS 어텐션 출력에 기록된 패치별 항을 분리하여 잠재적인 패치 레벨의 의미 증거를 복원합니다. 또한, 기여도를 기반으로 파생된 의미 응답을 의미-지오데식 토폴로지 게이트로 변환하여 밀집 특징 재구성을 위한 최종 레이어의 패치 유사성을 조정합니다. 진단 실험 결과, 이러한 기여도 특징은 강력한 로컬 의미 차별성과 텍스트 기반의 공간 정렬 능력을 보여줍니다. TraceCLIP은 추가적인 학습 없이, 외부 시각 기반 모델이나 영역 레벨 지도 없이, 다양한 백본 및 배경 설정을 포함하는 8개의 제로샷 의미 분할 벤치마크에서 기존의 가장 강력한 학습이 필요 없는 방법보다 평균 mIoU 점수가 1.3점에서 4.5점 향상된 성능을 달성했습니다. 더 넓은 관점에서, 이러한 결과는 공간적으로 위치화된 의미 정보가 전역적으로 정렬된 표현의 내부 구조 내에서 여전히 접근 가능할 수 있음을 시사합니다.

Original Abstract

Dense vision-language understanding, including object localization, region recognition, and open-vocabulary semantic segmentation, requires associating language concepts with spatially grounded visual regions. CLIP provides a strong foundation for these tasks by learning a shared image-text embedding space from large-scale contrastive pre-training. However, its image-level objective aligns text with a CLS-derived global representation, leaving local vision-language correspondence only indirectly constrained. Existing methods either introduce additional supervision, external models, or task-specific adaptation, while training-free approaches mainly recover dense responses from existing patch features without examining where local semantics become most accessible within CLIP. We introduce TraceCLIP, a training-free framework that recovers latent patch-level semantic evidence by isolating the patch-specific terms written into the CLS attention output. TraceCLIP further converts contribution-derived semantic responses into a semantic-geodesic topology gate that calibrates final-layer patch affinity for dense feature reconstruction. Diagnostic experiments show that these contribution features exhibit strong local semantic discrimination and text-conditioned spatial alignment. On eight zero-shot semantic segmentation benchmarks, TraceCLIP achieves gains of 1.3 to 4.5 points in average mIoU over the strongest prior training-free methods across both backbones and background settings, without additional training, external vision foundation models, or region-level supervision. More broadly, these findings suggest that spatially localized semantics may remain accessible within the internal construction of globally aligned representations.

0 Citations
0 Influential
6 Altmetric
30.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!