2605.25952v1 May 25, 2026 cs.CV

VEN-VL: 효과적이고 효율적인 다중 모드 이해를 위한 시각적 앙상블 Mixture of Experts 프레임워크

VEN-VL: A Visual Ensemble MoE Framework for Effective and Efficient Multi-Modal Understanding

Yujiu Yang
Yujiu Yang
Citations: 463
h-index: 8
Yinghao Wu
Yinghao Wu
Citations: 62
h-index: 4
Zhuoyan Luo
Zhuoyan Luo
Citations: 311
h-index: 6
Yiyao Yu
Yiyao Yu
Citations: 165
h-index: 6
Zhaojian Yu
Zhaojian Yu
Citations: 108
h-index: 5
Xiao-Ping Zhang
Xiao-Ping Zhang
Citations: 177
h-index: 6

최근의 효율적인 방법들이 다중 모드 이해 속도를 향상시키는 데 놀라운 발전을 이루었음에도 불구하고, 여전히 눈에 띄는 성능 저하가 발생합니다. 이러한 방법들은 단일 시각적 단서의 높은 압축률을 강조하고, 미세한 주의 정렬을 사용한 휴리스틱 기반 가지치기 전략에 의존하여 시각적 토큰의 정보 용량과 밀도를 제한하는 병목 현상을 야기합니다. 이러한 한계를 극복하기 위해, 우리는 '풍부하게 만든 후 압축한다(enrich then compact)' 원칙에 따라 효과적이고 효율적인 인식을 위한 시각적 앙상블 Mixture of Experts (MoE) 프레임워크인 VEN-VL을 제안합니다. 구체적으로, 우리는 먼저 다양한 관점의 시각적 표현을 통합하여 정보 용량을 풍부하게 만들고, 그 후 특수화된 시각적 전문가 내에서 적응형 라우터를 사용하여 정보를 점진적으로 압축하여 정보 밀도를 향상시킵니다. 또한, 명시적인 시각적 감독을 통해 기본적인 구조의 재구성 능력을 활용하여 중요한 정보 보존을 돕습니다. 실험 결과는 제한된 수의 정보가 응축된 토큰을 사용하는 복잡한 시각적 작업에서 우리의 우수성을 입증하며, 성능과 효율성 간의 격차를 효과적으로 해소합니다.

Original Abstract

Despite the remarkable progress achieved by recent efficient methods in accelerating multimodal understanding, they still suffer from noticeable performance degradation. Their emphasis on the high compression ratio of a single visual clue and reliance on the heuristic pruning strategy with coarse attention alignment incurs a bottleneck on the information capacity and density of visual tokens. Addressing this limitation, we propose VEN-VL, a visual ensemble MoE framework for effective and efficient perception following the enrich then compact principle. Specifically, we first enrich the information capacity by unifying the visual representations of different perspectives, and then progressively compact it with adaptive routers in specialized visual experts to enhance the information density. Furthermore, we incorporate the reconstruction ability of vanilla structure via explicit visual supervision, facilitating crucial information preservation. Experimental results demonstrate our superiority in complex visual tasks with few information-condensed tokens, which effectively bridges the gap between performance and efficiency.

0 Citations
0 Influential
4 Altmetric
20.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!