2603.25035v1 Mar 26, 2026 cs.AI

시각-언어 모델에서의 압축 현상에 대한 메커니즘적 해석

Mechanistically Interpreting Compression in Vision-Language Models

V. Elluru
V. Elluru
Citations: 21
h-index: 3
Arthur Singh
Arthur Singh
Citations: 1
h-index: 1
R. Aguero
R. Aguero
Citations: 155
h-index: 3
Ajayta Agarwal
Ajayta Agarwal
Citations: 0
h-index: 0
Debojyoti Das
Debojyoti Das
Citations: 12
h-index: 2
Hreetam Paul
Hreetam Paul
Citations: 0
h-index: 0

압축된 시각-언어 모델(VLMs)은 메모리 및 계산 비용을 줄이는 데 널리 사용되며, 실제 환경에 적용하기에 적합한 선택입니다. 그러나 이러한 모델을 압축하면 내부 연산 및 안전 기능이 유지되는지에 대한 우려가 제기됩니다. 본 연구에서는 인과적 회로 분석 및 크로스 코더 기반 특징 비교를 사용하여 가지치기(pruning) 및 양자화(quantization)가 대표적인 VLMs의 내부 구조를 어떻게 근본적으로 변화시키는지 조사했습니다. 연구 결과, 가지치기는 일반적으로 회로 구조를 유지하지만 내부 특징을 회전시키고 감쇠시키는 반면, 양자화는 회로를 더 높은 수준에서 수정하지만 생존하는 특징은 더 잘 정렬되는 것으로 나타났습니다. 이러한 통찰력을 바탕으로, 우리는 다양한 안전 범주에 걸쳐 유해한 입력과 일치하는 안전한 대조 예제를 결합한 새로운 벤치마크인 VLMSafe-420을 소개합니다. 연구 결과, 가지치기는 진정한 거부 동작이 급격하게 감소하는 것을 보여주며, 이는 압축 방식 선택이 안전에 영향을 미칠 수 있음을 시사합니다.

Original Abstract

Compressed vision-language models (VLMs) are widely used to reduce memory and compute costs, making them a suitable choice for real-world deployment. However, compressing these models raises concerns about whether internal computations and safety behaviors are preserved. In this work, we use causal circuit analysis and crosscoder-based feature comparisons to examine how pruning and quantization fundamentally change the internals across representative VLMs. We observe that pruning generally keeps circuit structure intact but rotates and attenuates internal features, while quantization modifies the circuits at a higher level yet leaves the surviving features better aligned. Leveraging this insight, we also introduce VLMSafe-420, a novel benchmark that pairs harmful inputs with matched benign counterfactuals across various safety categories. Our findings show that pruning causes a sharp drop in genuine refusal behavior, suggesting that the choice of compression has safety implications.

0 Citations
0 Influential
1.5 Altmetric
7.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!