2607.14711v1 Jul 16, 2026 cs.CV

VideoSEMA: 확장 가능하고 효율적인 Mamba 유사 어텐션 기반 비디오 이해 모델

VideoSEMA: a scalable and efficient Mamba-like attention for video understanding

N. Tran
N. Tran
Citations: 119
h-index: 3
Jiancheng Lyu
Jiancheng Lyu
Citations: 12
h-index: 1
Yunling Zheng
Yunling Zheng
Citations: 47
h-index: 4
Y. Qi
Y. Qi
Citations: 1,684
h-index: 18
Jack Xin
Jack Xin
Citations: 0
h-index: 0
Fanghui Xue
Fanghui Xue
Citations: 25
h-index: 4
Shuai Zhang
Shuai Zhang
Citations: 644
h-index: 7

본 논문에서는 비디오 이해 (분류)를 위한 분리된 시-공간 어텐션 모델인 VideoSEMA를 제안합니다. VideoSEMA는 공간 영역에서 확장 가능하고 효율적인 Mamba 유사 어텐션 블록 (SEMA)과 시간 영역에서 소프트맥스 어텐션을 사용합니다. 각 프레임에서 SEMA 어텐션은 Mamba의 거시 구조와 유사하게, 로컬 윈도우 어텐션을 글로벌 평균화와 병렬적으로 적용합니다. 특정 순위 조건 하에서, 계산 비용이 저렴한 분리된 시-공간 어텐션이 전체 시-공간 어텐션과 동등하다는 것을 증명합니다. K400 벤치마크 데이터 세트에서 VideoSEMA는 더 무거운 비전 트랜스포머 및 Mamba 모델보다 뛰어난 성능을 보입니다. SSv2 벤치마크 데이터에서도 VideoSEMA는 유사한 파라미터 크기의 다른 모델들 중에서 가장 높은 정확도를 달성합니다. K400 데이터 세트에서 이미지 해상도가 표준 $224^2$에서 $1024^2$로 증가할 때, 그리고 추가적인 튜닝 없이 VideoSEMA는 VideoMamba보다 훨씬 더 우수한 성능을 유지합니다. 확장된 시간 어텐션을 통해 VideoSEMA를 더 긴 비디오에 적용하는 것은 유망한 연구 방향입니다.

Original Abstract

We present for video understanding (classification) a split space-time attention model, VideoSEMA, consisting of a scalable and efficient Mamba-like attention (SEMA) block in space and a softmax temporal attention in time. In each frame, SEMA attention applies a local window attention in parallel with a global averaging in a Mamba macro-architecture, which is called Mamba-like. Under certain rank conditions, we prove that the computationally cheaper split space-time attention is equivalent to full space-time attention. On benchmark K400 data sets, VideoSEMA out-performs heavier vision transformer and Mamba models. On benchmark SSv2 data, VideoSEMA leads in top-1 accuracy among models of similar parameter sizes. As image resolution scales up from standard $224^2$ to $1024^2$ on K400 and without fine-tuning, VideoSEMA degrades much more gracefully than VideoMamba in accuracy. It is promising to extend VideoSEMA to longer videos with a dilated/sparse temporal attention.

0 Citations
0 Influential
9 Altmetric
45.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!