2604.20665v1 Apr 22, 2026 cs.CV

보는 데 드는 비용: 단일화된 패러다임 내에서 신뢰할 수 있는 다중 모드 추론을 달성하는 방법

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm

Karan Goyal
Karan Goyal
Citations: 15
h-index: 2
Dikshant Kukreja
Dikshant Kukreja
Citations: 1
h-index: 1

비전-언어 모델(VLMs)의 급속한 확산은 통합된 다중 모드 지식 탐색의 시작으로 널리 환영받고 있지만, 이러한 모델의 기반은 위험하고 의심받지 않는 전제, 즉 현재의 VLMs가 다중 모드 데이터를 충실하게 합성한다는 믿음에 기반합니다. 우리는 이에 반박하며, 지배적인 비전 인코더-프로젝터-LLM 패러다임에는 심각한 신뢰성 위기가 존재한다고 주장합니다. 최첨단 모델은 시각적 입력에서 의미 있는 지식을 추출하는 대신, 종종 기능적 시야 결여 현상을 보이며, 즉 강력한 언어적 선입견을 활용하여 심각한 시각적 표현상의 제약을 우회합니다. 본 연구에서는 다중 모드 평가의 기존 방법론에 도전합니다. 기존 방법은 데이터 제거 또는 새로운 데이터 세트 생성에 의존하며, 따라서 데이터 세트의 편향과 아키텍처의 한계를 치명적으로 혼동시킵니다. 우리는 정보 이론에 기반한 획기적인 접근 방식인 '모달리티 번역 프로토콜'을 제안합니다. 이를 통해 '보는 데 드는 비용'을 정량적으로 파악할 수 있습니다. 우리는 데이터를 제거하는 대신 의미 있는 정보를 번역함으로써, '보는 데 드는 비용'의 세 가지 새로운 지표인 '톨(ToS)', '저주(CoS)', '허위(FoS)'를 정의하고, 궁극적으로 '의미적 충분성 기준(SSC)'을 제시합니다. 또한, 우리는 다중 모드 확장의 '발산 법칙'이라는 도발적인 가설을 제시합니다. 즉, 기본 언어 엔진이 전례 없는 추론 능력을 갖추게 됨에 따라, 시각적 지식 병목 현상의 수학적 페널티가 역설적으로 증가한다는 것입니다. 우리는 KDD 커뮤니티에 '다중 모드 이득'이라는 환상적인 목표를 포기할 것을 촉구합니다. 우리는 SSC를 수동적인 진단 제약 조건에서 능동적인 아키텍처 설계 원칙으로 격상시켜, 차세대 AI 시스템이 실제로 데이터를 '보는' 진정한 다중 모드 추론을 달성할 수 있도록 하는 엄격하고 신뢰할 수 있는 기반을 제공합니다.

Original Abstract

The rapid proliferation of Vision-Language Models (VLMs) is widely celebrated as the dawn of unified multimodal knowledge discovery but its foundation operates on a dangerous, unquestioned axiom: that current VLMs faithfully synthesise multimodal data. We argue they do not. Instead, a profound crisis of trustworthiness underlies the dominant Vision Encoder-Projector-LLM paradigm. Rather than extracting grounded knowledge from visual inputs, state-of-the-art models frequently exhibit functional blindness, i.e., exploiting strong language priors to bypass severe visual representation bottlenecks. In this work, we challenge the conventional methodology of multimodal evaluation, which relies on data ablation or new dataset creation and therefore fatally conflates dataset biases with architectural incapacity. We propose a radical, information-theoretic departure: the Modality Translation Protocol, designed to quantifiably unmask the Expense of Seeing. By translating semantic payloads rather than ablating them, we formulate three novel metrics -- the Toll (ToS), Curse (CoS), and Fallacy (FoS) of Seeing -- culminating in the Semantic Sufficiency Criterion (SSC). Furthermore, we posit a provocative Divergence Law of Multimodal Scaling, hypothesising that as the underlying language engines scale to unprecedented reasoning capabilities, the mathematical penalty of the visual knowledge bottleneck paradoxically increases. We challenge the KDD community to abandon the illusory pursuit of "multimodal gain". By elevating the SSC from a passive diagnostic constraint to an active architectural blueprint, we provide the rigorous, trustworthy foundation required to force the next generation of AI systems to truly see the data, achieving true multimodal reasoning.

1 Citations
0 Influential
1 Altmetric
6.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!