2608.08832v1 Aug 09, 2026 cs.CV

Visual Token Codec: ViT 특징 코딩을 위한 공간 중복성 활용

Visual Token Codec: Unleashing Spatial Redundancy for ViT Feature Coding

Hongwei Hu
Hongwei Hu
Citations: 77
h-index: 4
Zhengxue Cheng
Zhengxue Cheng
Citations: 263
h-index: 9
Qi Wang
Qi Wang
Citations: 24
h-index: 3
Li Song
Li Song
Citations: 113
h-index: 6
Feng Zhang
Feng Zhang
Citations: 10
h-index: 1
Donghui Feng
Donghui Feng
Citations: 168
h-index: 6
Changsheng Gao
Changsheng Gao
Citations: 8
h-index: 2
Wenhan Yang
Wenhan Yang
Citations: 116
h-index: 5
Qunshan Gu
Qunshan Gu
Citations: 14
h-index: 2

대규모 비전 기반 모델의 분산 배포는 종종 ViT 백본을 분할하고 컴퓨팅 노드 간에 중간 토큰 특징을 교환하며, 이는 대역폭 및 계산 제약 조건 하에서 효율적인 특징 압축의 중요성을 강조합니다. 기존의 ViT 특징 코덱은 일반적으로 다양한 글로벌 및 패치 토큰을 L x C 크기의 유사 이미지로 평면화하여 엔트로피 모델이 주로 시퀀스 축 의존성만 포착하고 원본의 2차원 패치 그리드 구조를 간과하게 됩니다. 본 논문에서는 ViT 패치 토큰이 원래 그리드에서 강력한 국소적인 공간 상관관계를 유지한다는 것을 보여줍니다. 이러한 구조적 사전 지식을 활용하기 위해, 우리는 전역 토큰과 패치 토큰을 별도의 코딩 경로로 분리하는 이중 경로 기반 학습 코덱인 Visual Token Codec (VTC)를 제안합니다. 전역 토큰은 경량화된 요소 기반 사전(factorized prior)으로 압축되고, 패치 토큰은 공간-채널 컨텍스트 엔트로피 모델을 사용하여 패치-토큰 그리드에서 인코딩됩니다. 또한 VTC는 중간 레이어의 압축과 실용적인 비트율 적응을 지원하기 위해 후속 ViT 블록 이후에 특징 매칭 감독(feature-matching supervision)을 적용하고 단일 코덱 내부에 가변 비트율 모듈을 통합합니다. DINOv2 및 SAM3에 대한 실험 결과, VTC는 분류, 분할 및 탐지 작업에서 대표적인 ViT 특징 코딩 기준 성능을 지속적으로 능가하는 것으로 나타났습니다. 압축되지 않은 특징 성능의 90% 수준에서 VTC는 이러한 작업에서 비트율을 15.7배에서 37.4배까지 줄입니다. 또한 실용적인 전송 및 저장 중심 배포 시나리오를 위한 중간 레이어의 비트율-효율성 분석도 제공합니다.

Original Abstract

Distributed deployment of large vision foundation models often partitions a ViT backbone and exchanges intermediate token features between computing nodes, making efficient feature compression critical under bandwidth and computation constraints. Existing ViT feature codecs typically flatten heterogeneous global and patch tokens into an L x C pseudo image, causing entropy models to mainly capture sequence-axis dependencies while overlooking the native two-dimensional patch-grid structure. In this paper, we show that ViT patch tokens retain strong local spatial correlations on the original grid. To exploit this structural prior, we propose the Visual Token Codec (VTC), a dual-path learned codec that separates global and patch tokens into dedicated coding paths. Global tokens are compressed with a lightweight factorized prior, whereas patch tokens are encoded on the patch-token grid using a spatial-channel context entropy model. To support intermediate-layer compression and practical rate adaptation, VTC further incorporates feature-matching supervision after subsequent ViT blocks and variable-rate modules within a single codec. Experiments on DINOv2 and SAM3 show that VTC consistently outperforms representative ViT feature coding baselines on classification, segmentation, and detection tasks. At 90% of uncompressed-feature performance, VTC reduces bitrate by 15.7x-37.4x across these tasks. We further provide intermediate-layer rate-utility analyses for practical transmission- and storage-oriented deployment scenarios.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!