CRAFT: 비디오 토큰의 재귀적 적응적 융합을 통한 압축 - 시각-언어 모델을 위한 방법
CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models
비디오 이해에서, 시각-언어 모델(VLM)은 방대한 수의 시각 정보를 처리해야 하며, 이는 사전 학습 단계에서의 계산 및 메모리 비용을 크게 증가시킵니다. 이러한 시각 정보는 공간-시간 차원에서 높은 수준의 중복성을 갖지만, 높은 압축률은 종종 중요한 세부 정보 손실로 이어집니다. 기존의 토큰 압축 방법들은 효율성이 제한적이거나, 추가적인 모듈이 필요하여 비용이 많이 드는 정렬 학습을 요구하며, 효율성과 적응성 간의 균형을 맞추지 못하는 경우가 많습니다. 이러한 한계를 극복하기 위해, 우리는 비디오 토큰의 재귀적 적응적 융합을 통한 압축 방법인 CRAFT를 제안합니다. CRAFT는 파라미터가 없는 토큰 선택과 학습 가능한 토큰 융합을 분리하여 토큰을 재귀적으로 병합합니다. 전역적인 유사성은 어떤 토큰을 병합할지 결정하며, 위치 정보를 고려한 가중치 모듈과 콘텐츠에 적응하는 채널 기반 게이트는 어떻게 병합할지를 학습합니다. 전체 압축 파이프라인은 쿼리에 의존하지 않습니다. CRAFT는 유지되는 모든 토큰이 원래 토큰의 선형 조합이기 때문에, 원본 토큰의 실제 공간-시간 좌표를 보존하며 사전 학습된 언어 모델의 입력 분포와 일관성을 유지합니다. 다양한 비디오 벤치마크에서 수행한 실험 결과, CRAFT는 기존의 최첨단 토큰 압축 방법보다 우수한 성능을 지속적으로 보여줍니다. 약 $8 imes$ 수준의 압축률에서도 평균 정확도의 약 $97%$를 유지하며, 상당한 효율성 향상을 보입니다.
In video understanding, vision-language models (VLMs) must ingest massive numbers of visual tokens, causing the computational and memory cost of the prefill stage to rise sharply. Such visual sequences are highly redundant along the spatio-temporal dimension, yet a high compression ratio is often accompanied by the loss of critical details. Existing token-compression methods either employ heuristic, training-free compression with limited content adaptivity or introduce additional modules that require expensive alignment training, leaving the trade-off between efficiency and adaptivity unresolved. To alleviate this limitation, we propose CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens. CRAFT recursively merges tokens by decoupling parameter-free token selection from learnable token fusion: global similarity determines which tokens to merge, while a position-aware weighting module and a content-adaptive channel-wise gate learn how to fuse them. The whole compression pipeline is query-agnostic. Because every retained token is a linear combination of the original tokens, CRAFT preserves their true spatio-temporal coordinates and stays aligned with the pre-trained language model's input distribution. Experiments on multiple representative video benchmarks show that CRAFT consistently outperforms prior state-of-the-art token-compression methods. At about $8\times$ compression, it retains roughly $97\%$ of the backbone's average accuracy and shows significant efficiency improvement.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.