2608.03812v1 Aug 04, 2026 cs.CV

OmniPack: 효율적인 다중 모달 대규모 언어 모델을 위한 통합 토큰 압축

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Haotian Wang
Haotian Wang
Citations: 75
h-index: 5
Liang Ding
Liang Ding
University of Sydney / JD Explore Academy
Citations: 5,343
h-index: 40
Chengfu Huo
Chengfu Huo
Citations: 264
h-index: 6
Wanshun Su
Wanshun Su
Citations: 15
h-index: 2
Yan Min
Yan Min
Citations: 47
h-index: 3
Peng Wu
Peng Wu
Citations: 0
h-index: 0

다중 모달 대규모 언어 모델(Omni-LLM)은 오디오-비주얼 이해 작업에서 뛰어난 성능을 보이고 있지만, 긴 시퀀스의 중복된 오디오 및 비주얼 토큰을 처리하는 데 상당한 계산 비용이 발생하며, 효율적인 배포를 위해서는 적극적인 토큰 압축이 필요합니다. 기존 방법들은 종종 낮은 토큰 예산 환경에서 성능 저하를 겪습니다. LLM 이전 단계의 압축은 구조적으로 중요한 정보를 삭제하거나 전역적으로 분산된 증거를 손실시키는 반면, LLM 내부의 압축은 종종 쿼리 기반 오디오-비주얼 협력을 충분히 활용하지 못합니다. 이러한 한계를 해결하기 위해, 우리는 LLM 이전 단계에서 구조적 압축을 수행하고, 동시에 LLM 내부에서 작업 관련 의미론적 개선을 통합하는 학습이 필요 없는 프레임워크인 OmniPack을 제안합니다. OmniPack은 LLM 전에 모달리티별 중요도, 전역 커버리지 및 유사성 기반 병합을 통해 구조적 중복성을 제거합니다. 충분한 다중 모달 상호 작용 후에는 텍스트 가이드 및 오디오-비주얼 협력을 통해 다양한 작업 관련 표현을 더욱 통합합니다. 세 가지 Omni-LLM 백본 모델을 사용한 다섯 가지 벤치마크에서의 광범위한 실험 결과, OmniPack은 다양한 유지 비율에서 일관되게 최고의 성능-효율성 균형을 달성하며 기존 방법들보다 우수한 성능을 보였습니다. 특히, Qwen2.5-Omni-7B 모델에서 OmniPack은 원래 성능의 98.0%를 유지하면서 FLOPs를 16.7%로 줄였으며, 원래 FLOPs의 6.8%만으로도 원래 성능의 92.9%를 유지할 수 있었습니다.

Original Abstract

Omni-modal large language models (Omni-LLMs) have achieved remarkable performance on audio-visual understanding tasks, but processing long and highly redundant visual and audio token sequences incurs substantial computational overhead, demanding aggressive token compression for efficient deployment. Existing methods often degrade at low token budgets: pre-LLM compression may discard structurally important and globally distributed evidence, whereas inner-LLM compression often underexploits query-conditioned audio-visual collaboration. To address these limitations, we propose OmniPack, a training-free framework that coordinates structural compression before the LLM with task-relevant semantic refinement within the LLM. Before the LLM, OmniPack removes structural redundancy through modality-specific importance, global coverage, and similarity-aware merging. After sufficient multimodal interaction, it further consolidates diverse, task-relevant representations through textual guidance and audio-visual collaboration. Extensive experiments on five benchmarks with three Omni-LLM backbones demonstrate that OmniPack consistently achieves the best performance-efficiency trade-off across diverse retention ratios, outperforming all existing methods. Notably, on Qwen2.5-Omni-7B, OmniPack preserves 98.0% of the original performance while reducing FLOPs to 16.7%, and still retains 92.9% of the original performance with only 6.8% of the original FLOPs.

0 Citations
0 Influential
20 Altmetric
100.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!