2607.25527v1 Jul 28, 2026 cs.CV

Argus-Unified: 이미지 이해 및 생성을 위한 효율적이고 경제적인 통합 모델 개발

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation

Jingtao Li
Jingtao Li
Citations: 78
h-index: 5
Weiming Zhuang
Weiming Zhuang
Citations: 1,038
h-index: 15
Chen Chen
Chen Chen
Citations: 342
h-index: 10
Sina Sajadmanesh
Sina Sajadmanesh
Idiap Research Institute, EPFL
Citations: 729
h-index: 9
Lingjuan Lyu
Lingjuan Lyu
Citations: 62
h-index: 4
Jiabo Huang
Jiabo Huang
Citations: 284
h-index: 8
Z. Li
Z. Li
Citations: 18
h-index: 2

시각적 이해와 생성 기능을 하나의 모델로 통합하는 것은 엄청난 잠재력을 지니지만, 막대한 계산량과 데이터 요구 사항, 그리고 두 기능에 필요한 시각적 특징 간의 충돌 때문에 어려운 과제였습니다. 이러한 문제점을 해결하기 위해, 우리는 낮은 계산 및 데이터 요구량을 갖춘 효율적인 통합 멀티모달 모델인 Argus-Unified를 제안합니다. Argus-Unified는 처음부터 모달리티 정렬을 수행하는 대신, 강력한 멀티모달 사전 지식을 제공하는 사전 학습된 비전-언어 모델(VLMs)을 효과적으로 활용합니다. 특히, 우리는 연속적인 토큰은 이해에 사용하고, 불연속적인 토큰은 생성에 사용하도록 설계된 하이브리드 시각적 토큰을 도입하여, 통합된 비전 인코더를 고정 상태로 유지하면서 학습을 진행합니다. 우리의 훈련 파이프라인은 두 단계로 구성됩니다. 첫 번째 단계에서는 고정된 비전 인코더 위에 양자화기(quantizer)와 이미지 디코더를 학습하고, 두 번째 단계에서는 사전 학습된 VLM에서 초기화된 LLM을 통합 멀티모달 모델링을 위해 훈련합니다. 우리는 이전 연구보다 훨씬 적은 데이터(1560만 개)와 가장 낮은 비용(~2,000 달러)으로 훈련하여, 통합 멀티모달 모델이 강력한 성능을 유지하면서도 경제적으로 훈련될 수 있음을 입증했습니다. 주목할 만한 점은, 우리의 모델이 GQA, POPE 및 VQAv2에서 최첨단 수준의 멀티모달 이해 능력을 달성했으며, 전용 비전 인코더를 사용하는 모델(예: Janus, Janus-Pro)과 비교하여 경쟁력 있는 생성 품질을 보이며, 동시에 약 10배 낮은 비용과 약 5배 적은 데이터로 이를 달성했습니다. 우리는 Argus-Unified가 통합 모델 개발의 진입 장벽을 낮추는 유용한 기준 모델이 될 수 있을 것으로 기대합니다.

Original Abstract

Unifying visual understanding and generation in one model holds immense promise, but remains challenging and expensive due to heavy compute and data demands and conflicts between the visual features needed for these two capabilities. To address these challenges, we present Argus-Unified, a compact, effective and unified multimodal model built with low demand on computation and data. Instead of aligning modalities from scratch, Argus-Unified effectively leverages pretrained vision-language models (VLMs) that provide strong multimodal priors. Specifically, we introduce hybrid visual tokens that preserve continuous tokens for understanding while learning discrete tokens for generation from a frozen unified vision encoder. Our training pipeline includes two stages: the first stage learns a quantizer and image decoder on top of the frozen vision encoder, the second stage trains the LLM initialized from a pretrained VLM for the unified multimodal modeling. Using by far the least amount of data (15.6M) and the lowest cost (~$2,000), we demonstrate that unified multimodal models can be trained economically while achieving strong performance in both understanding and generation. Notably, our model attains state-of-the-art multimodal understanding on GQA, POPE, and VQAv2, and competitive generation quality compared to models with dedicated vision encoders (e.g., Janus, Janus-Pro), all at ~10x lower cost and with ~5x less data. We envision Argus-Unified as a useful baseline that lowers the development barrier for unified models.

0 Citations
0 Influential
7.5 Altmetric
37.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!