2602.07550v1 Feb 07, 2026 cs.CV

DINOv3의 의미 선택 격차를 훈련 불필요한 소량 샘플 분할을 통해 밝히기

Revealing the Semantic Selection Gap in DINOv3 through Training-Free Few-Shot Segmentation

Hussni Mohd Zakir
Hussni Mohd Zakir
Citations: 1
h-index: 1
E. Ho
E. Ho
Citations: 1
h-index: 1

최근의 자기 지도 학습 기반 비전 트랜스포머(ViT) 모델, 특히 DINOv3는 다양한 시각적 작업에 유용한 풍부한 특징 표현을 제공합니다. 본 연구에서는 클래스별 프로토타입과 그램 행렬 정제를 활용하는 훈련 불필요한 기준 모델인 FSSDINO를 통해, 고정된 DINOv3 특징의 고유한 소량 샘플 의미 분할(FSS) 능력을 조사합니다. 이진, 다중 클래스, 그리고 교차 도메인(CDFSS) 벤치마크에서의 결과는, 최종 백본 레이어에 적용된 이 간단한 접근 방식이 복잡한 디코더 또는 테스트 시간 적응을 포함하는 특수화된 방법과 경쟁력이 높다는 것을 보여줍니다. 더욱 중요하게는, 오라클(Oracle) 기반 레이어 분석을 통해 표준 마지막 레이어 특징과 전역적으로 최적의 중간 표현 사이의 상당한 성능 격차가 있음을 확인했습니다. 우리는 "안전 vs. 최적"이라는 딜레마를 밝혀냈습니다. 오라클은 더 높은 성능이 가능하며, 계산 집약적인 적응 방법의 결과를 따라잡을 수 있지만, 현재의 비지도 학습 및 지원 기반 선택 메트릭은 지속적으로 마지막 레이어 기준 성능보다 낮은 성능을 보입니다. 이는 기초 모델에서 발생하는 "의미 선택 격차"를 나타내며, 기존의 휴리스틱이 고품질 특징을 안정적으로 식별하는 데 실패하는 현상입니다. 본 연구는 "마지막 레이어"가 강력한 기준 성능을 보이는 것처럼 보이도록 하며, DINOv3의 잠재적인 의미 정보를 엄격하게 진단합니다. 코드는 https://github.com/hussni0997/fssdino 에서 공개적으로 이용할 수 있습니다.

Original Abstract

Recent self-supervised Vision Transformers (ViTs), such as DINOv3, provide rich feature representations for dense vision tasks. This study investigates the intrinsic few-shot semantic segmentation (FSS) capabilities of frozen DINOv3 features through a training-free baseline, FSSDINO, utilizing class-specific prototypes and Gram-matrix refinement. Our results across binary, multi-class, and cross-domain (CDFSS) benchmarks demonstrate that this minimal approach, applied to the final backbone layer, is highly competitive with specialized methods involving complex decoders or test-time adaptation. Crucially, we conduct an Oracle-guided layer analysis, identifying a significant performance gap between the standard last-layer features and globally optimal intermediate representations. We reveal a "Safest vs. Optimal" dilemma: while the Oracle proves higher performance is attainable, matching the results of compute-intensive adaptation methods, current unsupervised and support-guided selection metrics consistently yield lower performance than the last-layer baseline. This characterizes a "Semantic Selection Gap" in Foundation Models, a disconnect where traditional heuristics fail to reliably identify high-fidelity features. Our work establishes the "Last-Layer" as a deceptively strong baseline and provides a rigorous diagnostic of the latent semantic potentials in DINOv3.The code is publicly available at https://github.com/hussni0997/fssdino.

2 Citations
2 Influential
29.45879734614 Altmetric
11.0 Score
Original PDF
5

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!