통합 다중 모달 모델은 단일 공간에서 사고하는가? 교차 분기 제어(Cross-Branch Steering)를 통한 탐구
Do Unified Multimodal Models Think in One Space? A Lens Through Cross-Branch Steering
통합 다중 모달 모델(Unified Multimodal Models, UMMs)은 이해와 생성 능력을 하나의 아키텍처 내에 통합하는 것을 목표로 하지만, 이러한 기능들이 단일하고 전이 가능한 의미 공간을 공유하는지 여부는 아직 명확하지 않습니다. 이 질문은 근본적으로 어렵습니다. 왜냐하면 두 가지 분기는 이질적인 표현(텍스트 토큰 vs.\ 시각적 잠재 벡터)과 서로 다른 학습 목표를 가지고 있기 때문에 직접적인 비교가 어렵기 때문입니다. 이를 해결하기 위해, 우리는 하나의 분기로부터 의미 방향을 추출하여 다른 분기에 적용하는 개입 기반 프레임워크인 \"교차 분기 의미 제어(cross-branch semantic steering)\"를 소개합니다. 우리는 이해 분기에서 학습된 제어 벡터가 생성에 전이될 수 있으며, 이를 통해 제어 가능한 이미지 합성이 가능하고 의미적 정확도가 향상됨을 보여줍니다. 반대로, 역방향은 일관되게 제한적인 효과만을 나타냅니다. 우리의 분석 결과는 이러한 비대칭성이 실질적인 표현 불일치와 관련되어 있을 가능성을 시사합니다. 즉, 이해에서 파생된 벡터는 전이 가능한 객체 중심의 의미를 포착하는 반면, 생성에서 파생된 벡터는 주로 저수준 외관 특징을 인코딩합니다. 우리의 결과는 아키텍처 통합이 반드시 의미적 정렬을 보장하지 않으며, 교차 분기 제어가 다중 모달 표현을 탐구하기 위한 실용적인 도구임을 보여줍니다.
Unified multimodal models (UMMs) aim to integrate understanding and generation within a single architecture, yet it remains unclear whether these capabilities share a unified and transferable semantic space. This question is fundamentally challenging, as the two branches operate over heterogeneous representations (text tokens vs.\ visual latents) and distinct training objectives, making direct comparison difficult. To address this, we introduce \emph{cross-branch semantic steering}, an intervention-based framework that extracts semantic directions from one branch and applies them to the other. We show that steering vectors learned from the understanding branch can transfer to generation, enabling controllable image synthesis and improved semantic faithfulness. In contrast, the reverse direction consistently shows limited effectiveness. Our analysis suggests that this asymmetry may be related to a practical representational mismatch: understanding-derived vectors capture transferable, object-centric semantics, while generation-derived vectors primarily encode low-level appearance features. Our results reveal that architectural unification does not guarantee semantic alignment, and establish cross-branch steering as a practical tool for probing multimodal representations.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.