VideoSEMA: 확장 가능하고 효율적인 Mamba 유사 어텐션 기반 비디오 이해 모델
VideoSEMA: a scalable and efficient Mamba-like attention for video understanding
본 논문에서는 비디오 이해 (분류)를 위한 분리된 시-공간 어텐션 모델인 VideoSEMA를 제안합니다. VideoSEMA는 공간 영역에서 확장 가능하고 효율적인 Mamba 유사 어텐션 블록 (SEMA)과 시간 영역에서 소프트맥스 어텐션을 사용합니다. 각 프레임에서 SEMA 어텐션은 Mamba의 거시 구조와 유사하게, 로컬 윈도우 어텐션을 글로벌 평균화와 병렬적으로 적용합니다. 특정 순위 조건 하에서, 계산 비용이 저렴한 분리된 시-공간 어텐션이 전체 시-공간 어텐션과 동등하다는 것을 증명합니다. K400 벤치마크 데이터 세트에서 VideoSEMA는 더 무거운 비전 트랜스포머 및 Mamba 모델보다 뛰어난 성능을 보입니다. SSv2 벤치마크 데이터에서도 VideoSEMA는 유사한 파라미터 크기의 다른 모델들 중에서 가장 높은 정확도를 달성합니다. K400 데이터 세트에서 이미지 해상도가 표준 $224^2$에서 $1024^2$로 증가할 때, 그리고 추가적인 튜닝 없이 VideoSEMA는 VideoMamba보다 훨씬 더 우수한 성능을 유지합니다. 확장된 시간 어텐션을 통해 VideoSEMA를 더 긴 비디오에 적용하는 것은 유망한 연구 방향입니다.
We present for video understanding (classification) a split space-time attention model, VideoSEMA, consisting of a scalable and efficient Mamba-like attention (SEMA) block in space and a softmax temporal attention in time. In each frame, SEMA attention applies a local window attention in parallel with a global averaging in a Mamba macro-architecture, which is called Mamba-like. Under certain rank conditions, we prove that the computationally cheaper split space-time attention is equivalent to full space-time attention. On benchmark K400 data sets, VideoSEMA out-performs heavier vision transformer and Mamba models. On benchmark SSv2 data, VideoSEMA leads in top-1 accuracy among models of similar parameter sizes. As image resolution scales up from standard $224^2$ to $1024^2$ on K400 and without fine-tuning, VideoSEMA degrades much more gracefully than VideoMamba in accuracy. It is promising to extend VideoSEMA to longer videos with a dilated/sparse temporal attention.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.