JetViT: 고해상도 이미지에 대한 효율적인 비전 트랜스포머 - 학습 후 어텐션 검색 기반
JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search
본 논문에서는 JetViT라는 새로운 하이브리드 아키텍처를 가진 비전 트랜스포머(ViT) 모델 패밀리를 소개합니다. JetViT는 최첨단 풀 어텐션 기반 비전 모델의 정확도와 동등한 성능을 제공하면서, 고해상도 이미지에 대한 추론 효율성을 크게 향상시킵니다. 저희 접근 방식의 핵심은 Post-Training Attention Search라는 학습 후 어텐션 검색 프레임워크로, 이 프레임워크는 사전 훈련된 풀 어텐션 ViT 모델을 효율적인 하이브리드 어텐션 변형으로 변환합니다. 이는 불필요한 풀 어텐션 블록을 식별하고 선형 또는 윈도우 어텐션 블록으로 대체하는 방식으로 작동합니다. Post-Training Attention Search는 기본 모델의 MLP 및 어텐션 가중치를 상속하여 세 가지 주요 단계를 통해 효율적으로 아키텍처 설계 공간을 탐색합니다. (1) 선형 어텐션 블록 디자인 최적화, (2) 선형 어텐션 및 윈도우 어텐션 블록의 최적 조합 탐색, (3) 중요한 풀 어텐션 블록 식별 및 유지입니다. 저희는 JetViT를 두 가지 대표적인 고해상도 비전 모델인 DINOv3와 DepthAnythingV2에 대해 평가했습니다. NVIDIA H100 GPU에서 JetViT는 정확도를 손실하지 않고 최대 1.79배 더 높은 처리량과 최대 44.81% 더 낮은 지연 시간을 달성합니다. 저희는 곧 저희의 코드 및 가속화된 ViT 모델을 공개할 예정입니다.
We introduce JetViT, a novel family of hybrid-architecture Vision Transformer (ViT) models that match the accuracy of state-of-the-art full-attention vision foundation models while achieving substantially higher inference efficiency on high-resolution images. At the core of our approach is Post-Training Attention Search, a post-training acceleration framework that converts pre-trained full-attention ViTs into efficient hybrid-attention variants by identifying and replacing redundant full-attention blocks with linear or window-attention blocks. By inheriting the MLP and attention weights from the base model, Post-Training Attention Search efficiently explores the architectural design space through three key steps: (1) optimizing the linear-attention block design; (2) finding the best combination of linear-attention and window-attention blocks; and (3) identifying and preserving critical full-attention blocks. We evaluate JetViT on two representative high-resolution vision foundation models, DINOv3 and DepthAnythingV2. On the NVIDIA H100 GPU, JetViT achieves up to 1.79x higher throughput and up to 44.81% lower latency without sacrificing accuracy. We will release our code and accelerated ViT models soon.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.