2604.10949v1 Apr 13, 2026 cs.CV

가짜 통합: 엔트로피 탐색이 밝혀낸 통합 멀티모달 모델의 상이한 정보 패턴

Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models

Songlin Yang
Songlin Yang
Citations: 129
h-index: 3
Xianghao Kong
Xianghao Kong
HKUST
Citations: 330
h-index: 6
A. Rao
A. Rao
Citations: 51
h-index: 4

통합 멀티모달 모델(UMM)은 대규모 언어 모델(LLM)의 추론 능력과 비전 모델의 생성 능력을 결합하도록 설계되었습니다. 그러나 실제로 이러한 시너지 효과는 여전히 희미합니다. UMM은 LLM과 유사한 추론 능력을 이미지 생성에 적용하지 못하며, 상이한 응답 행동을 보입니다. 우리는 이러한 현상을 '가짜 통합'이라고 부릅니다. 이러한 현상의 내부 원인을 진단하는 것은 중요하지만, 기존의 탐색 방법은 모델 내부 정보 부족하거나 프롬프트-응답 의존성을 간과합니다. 이러한 한계를 해결하기 위해, 우리는 UMM이 입력을 인코딩하고 출력을 생성하는 방식을 동시에 분석하는 정보 이론 기반의 탐색 프레임워크를 제안합니다. 10개의 대표적인 UMM에 이 프레임워크를 적용한 결과, 가짜 통합은 다음과 같은 이중적 차이에서 비롯된다는 것을 알 수 있습니다. (i) 모달성 비대칭 인코딩: 비전과 언어가 서로 다른 엔트로피 경로를 따르는 현상, 그리고 (ii) 패턴 분리 응답: 텍스트 생성은 높은 엔트로피의 창의성을 보이는 반면, 이미지 생성은 낮은 엔트로피의 충실성을 강조하는 현상입니다. 컨텍스트 예측과 같이 양쪽을 통합하는 모델만이 더 진정한 통합을 달성하며, 더 적은 파라미터로도 강력한 추론 기반의 텍스트-이미지 생성이 가능합니다. 본 연구는 통합에 대한 최초의 모델 내부 탐색을 제공하며, 진정한 멀티모달 시너지는 공유된 파라미터뿐만 아니라 정보 흐름의 일관성이 필요하다는 것을 보여줍니다.

Original Abstract

Unified multimodal models (UMMs) were designed to combine the reasoning ability of large language models (LLMs) with the generation capability of vision models. In practice, however, this synergy remains elusive: UMMs fail to transfer LLM-like reasoning to image synthesis and exhibit divergent response behaviors. We term this phenomenon pseudo-unification. Diagnosing its internal causes is important, but existing probing methods either lack model-internal insight or ignore prompt-response dependencies. To address these limitations, we propose an information-theoretic probing framework that jointly analyzes how UMMs encode inputs and generate outputs. Applied to ten representative UMMs, our framework reveals that pseudo-unification stems from a dual divergence: (i) Modality-Asymmetric Encoding, where vision and language follow different entropy trajectories, and (ii) Pattern-Split Response, where text generation exhibits high-entropy creativity while image synthesis enforces low-entropy fidelity. Only models that unify both sides (e.g., via contextual prediction) achieve more genuine unification, enabling stronger reasoning-based text-to-image generation even with fewer parameters. Our work provides the first model-internal probing of unification, demonstrating that real multimodal synergy requires consistency in information flow, not just shared parameters.

1 Citations
0 Influential
3 Altmetric
16.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!