2606.24716v1 Jun 23, 2026 cs.CV

개념 어노테이션을 활용한 희소 오토인코더의 해석 가능성 평가

Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations

Jonas Klotz
Jonas Klotz
Citations: 6
h-index: 2
C. Dantas
C. Dantas
Citations: 303
h-index: 11
P. Jain
P. Jain
Citations: 1,964
h-index: 23
Diego Marcos
Diego Marcos
Citations: 109
h-index: 6
B. Demir
B. Demir
Citations: 320
h-index: 10

희소 오토인코더(SAE)는 시각 및 시각-언어 모델에서 해석 가능한 개념을 추출하는 데 점점 더 많이 사용되고 있지만, 기존의 평가 방법은 주로 프록시 메트릭이나 질적 검사에 의존하며 의미론적 일관성을 측정하지 못합니다. 본 연구에서는 사용자 연구 없이 SAE 잠재 변수와 인간이 어노테이션한 개념 간의 정렬을 정량화하는 인간 기반 평가 프레임워크를 제시하고, 이를 목표 속성 변화(attribute perturbations)를 통해 검증합니다. 시각 분야에서 이러한 개입형 평가를 가능하게 하기 위해, 단 하나의 속성이 다르게 변경된 이미지 쌍으로 구성된 합성 벤치마크인 synCUB 및 synCOCO를 구축했습니다. 본 연구에서는 SAE 잠재 변수와 어노테이션된 개념 간의 다대일 매핑을 지원하는 연합 기반 매칭 절차인 Fully-Binary Matching Pursuit (FBMP)를 소개하며, 이는 기존의 일대일 방법보다 우수한 성능을 보입니다. 기능적 검증을 위해, 목표 속성 변화에 따른 반응을 선택적으로 측정하고 예상되는 방향으로 반응하는지 확인하는 Targeted Attribute Perturbation Alignment Score (TAPAScore)를 제안합니다. 안전성 검사를 통해, 본 연구에서 제시한 매칭 및 TAPAScore는 훈련된 SAE와 훈련되지 않은 SAE를 안정적으로 구별하는 유일한 평가 지표임이 밝혀졌습니다. CLIP 및 DINOv2 임베딩으로 훈련된 다양한 SAE 모델을 분석한 결과, 과도한 오버컴플리트함은 속성 정렬을 감소시켜 해석 가능성을 저하시킬 수 있음을 확인했습니다. 본 연구의 평가 프레임워크는 적당한 크기의 사전(dictionary)이 가장 좋은 성능과 해석 가능성의 균형을 제공함을 시사합니다. 코드 및 데이터셋은 https://github.com/JonasKlotz/sae-concept-eval 에서 확인할 수 있습니다.

Original Abstract

Sparse autoencoders (SAEs) are increasingly used to extract interpretable concepts from vision and vision language models, yet existing evaluation methods largely rely on proxy metrics or qualitative inspection rather than measuring semantic correspondence. We present a human-grounded evaluation framework that quantifies alignment between SAE latents and human-annotated concepts, without requiring user studies, and validate this matching through targeted attribute perturbations. To enable this intervention-style evaluation in vision, we construct synCUB and synCOCO, synthetic benchmarks of paired images that differ in exactly one attribute. We introduce Fully-Binary Matching Pursuit (FBMP), a coalition-based matching procedure that supports many-to-one mappings between SAE latents and annotated concepts, and consistently outperforms one-to-one baselines. For functional validation, we propose a Targeted Attribute Perturbation Alignment Score (TAPAScore), which tests whether matched concepts respond selectively and in the expected direction under targeted image-level attribute perturbations. Under sanity checks, our matching and TAPAScore are the only evaluated metrics that reliably distinguish trained SAEs from untrained ones. Across SAEs trained on CLIP and DINOv2 embeddings, we find that increased overcompleteness can reduce perturbation alignment, indicating a reduction in interpretability. Our evaluation framework suggests that moderate dictionary sizes provide the best trade-off, yielding the most interpretable SAEs. Code and datasets are available at https://github.com/JonasKlotz/sae-concept-eval.

0 Citations
0 Influential
34.9657359028 Altmetric
0.0 Score
Original PDF
1

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!