2606.19932v1 Jun 18, 2026 cs.CV

공간 인지 감소 프레임워크: 효율적이고 정확한 시각 상태 공간 모델을 향하여

Spatial-Aware Reduction Framework: Towards Efficient and Faithful Visual State Space Models

Xiaofeng Wang
Xiaofeng Wang
Citations: 1,791
h-index: 17
Jindi Lv
Jindi Lv
Citations: 72
h-index: 5
Zheng Zhu
Zheng Zhu
Citations: 1,499
h-index: 17
Jiancheng Lv
Jiancheng Lv
Citations: 17
h-index: 2
Yuhao Zhou
Yuhao Zhou
Citations: 477
h-index: 9
Qing Ye
Qing Ye
Citations: 346
h-index: 8
Aoyu Li
Aoyu Li
Citations: 112
h-index: 4
Yueqi Duan
Yueqi Duan
Citations: 12
h-index: 2
Wentao Feng
Wentao Feng
Citations: 105
h-index: 2

Mamba는 긴 시각 정보를 모델링하는 데 뛰어난 효율성을 보여줍니다. 그러나 구조적으로 강화된 Mamba 변형에 토큰 감소를 적용할 때, 이러한 모델들은 심각한 성능 저하를 보입니다. 우리는 이러한 성능 저하의 원인을 기존의 감소 방법들이 가진 공간적 무관심성에서 찾습니다. 이는 선택적 스캔 메커니즘이 요구하는 2차원 구조적 전제를 위반하기 때문입니다. 본 연구에서는, 압축 과정 전반에 걸쳐 구조적 완전성을 유지하도록 설계된 공간 인지 토큰 감소 프레임워크인 STORM을 제안합니다. STORM은 감소를 그리드 토폴로지와 이웃 관계 일관성을 유지하는 로컬 제약을 적용한 공간 단위에서의 정형화된 연산으로 재정의합니다. STORM은 별도의 학습 없이 기존의 감소 파이프라인에 쉽게 통합될 수 있는 플러그 앤 플레이 모듈입니다. 실험 결과는, STORM이 다양한 시각 Mamba 모델에서 훈련 과정 없이도 최첨단 수준의 가지치기 정확도를 달성함을 보여줍니다. 특히, STORM은 VMamba에서 상당한 정확도 회복을 제공하며, 최고 63.3%의 top-1 정확도 향상을 보였습니다. 반면, STORM은 PlainMamba에서 1.0%의 정확도 감소만 발생시켜 ViT와 비교 가능한 성능을 달성했습니다.

Original Abstract

Mamba demonstrates strong efficiency in modeling long visual sequences. However, when token reduction is applied to structurally enhanced Mamba variants, these models exhibit a severe performance collapse. We attribute this degradation to the spatially agnostic nature of existing reduction methods, which violate the two-dimensional structural premise required by the selective scanning mechanism. In this work, we propose STORM, a spatial-aware token reduction framework designed to maintain structural integrity throughout the compression process. STORM reformulates reduction into a structured operation on spatial units, enforcing localized constraints to maintain both grid topology and neighborhood coherence. As a plug-and-play module, STORM equips existing reduction pipelines with explicit spatial awareness without any training. Empirical results demonstrate that STORM achieves state-of-the-art pruning accuracy across diverse vision Mamba backbones under training-free settings. Notably, STORM delivers a substantial accuracy recovery on VMamba, outperforming prior methods by up to 63.3\% in top-1 accuracy. Meanwhile, STORM incurs only a 1.0\% accuracy drop on PlainMamba, achieving performance comparable to ViT.

0 Citations
0 Influential
8.5 Altmetric
42.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!