자기 지도 학습 기반 비전 트랜스포머 모델에서의 인간과 유사한 객체 그룹화
Human-like Object Grouping in Self-supervised Vision Transformers
자기 지도 학습 방식으로 훈련된 비전 기반 모델은 다양한 작업에서 뛰어난 성능을 보이며, 객체 분할과 관련된 새로운 기능을 보여줍니다. 그러나 이러한 모델들이 인간의 객체 인식과 얼마나 일치하는지는 아직 명확히 밝혀지지 않았습니다. 본 연구에서는 자연스러운 장면에서 제시되는 점 쌍에 대해 참가자들이 동일/다른 객체에 대한 판단을 내리는 행동 기반 벤치마크를 도입했습니다. 이는 고전적인 심리학 실험 방식을 1000회 이상의 실험으로 확장한 것입니다. 다양한 비전 모델을 사용하여 모델의 표현을 기반으로 참가자들의 반응 시간을 예측하고, 모델 세대별로 꾸준한 성능 향상을 확인했습니다. 모델 구조와 학습 목표 모두 이러한 일치도 향상에 기여하며, 특히 DINO 자기 지도 학습 방식으로 훈련된 트랜스포머 기반 모델이 가장 뛰어난 성능을 보였습니다. 이러한 개선의 원인을 분석하기 위해, 이미지 패치 간의 유사성을 측정하여 표현의 객체 중심적인 요소를 정량화하는 새로운 지표를 제안했습니다. 모델 전반적으로, 객체 중심적인 구조가 더 강할수록 인간의 분할 행동을 더 정확하게 예측하는 것을 확인했습니다. 또한, 지도 학습 기반 트랜스포머 모델의 Gram 행렬을 자기 지도 학습 모델의 Gram 행렬과 증류(distillation)를 통해 일치시키면, 인간의 행동과 더 잘 일치하게 된다는 것을 보여주었습니다. 이는 이전 연구에서 Gram 행렬 앵커링이 DINOv3의 특징 품질을 향상시킨다는 결과를 뒷받침합니다. 종합적으로, 본 연구 결과는 자기 지도 학습 기반 비전 모델이 객체 구조를 인간과 유사한 방식으로 학습하며, Gram 행렬의 구조가 지각적인 일치도를 높이는 데 중요한 역할을 한다는 것을 시사합니다.
Vision foundation models trained with self-supervised objectives achieve strong performance across diverse tasks and exhibit emergent object segmentation properties. However, their alignment with human object perception remains poorly understood. Here, we introduce a behavioral benchmark in which participants make same/different object judgments for dot pairs on naturalistic scenes, scaling up a classical psychophysics paradigm to over 1000 trials. We test a diverse set of vision models using a simple readout from their representations to predict subjects' reaction times. We observe a steady improvement across model generations, with both architecture and training objective contributing to alignment, and transformer-based models trained with the DINO self-supervised objective showing the strongest performance. To investigate the source of this improvement, we propose a novel metric to quantify the object-centric component of representations by measuring patch similarity within and between objects. Across models, stronger object-centric structure predicts human segmentation behavior more accurately. We further show that matching the Gram matrix of supervised transformer models, capturing similarity structure across image patches, with that of a self-supervised model through distillation improves their alignment with human behavior, converging with the prior finding that Gram anchoring improves DINOv3's feature quality. Together, these results demonstrate that self-supervised vision models capture object structure in a behaviorally human-like manner, and that Gram matrix structure plays a role in driving perceptual alignment.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.