2608.05381v1 Aug 05, 2026 cs.AI

C$^3$PO: 다중 모드 모델의 교차 모드 통합 및 반사실적 성능 평가

C$^3$PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models

P. Kumaraguru
P. Kumaraguru
Citations: 11,839
h-index: 46
Tanuja Ganu
Tanuja Ganu
Citations: 982
h-index: 13
Swapnanil Mukherjee
Swapnanil Mukherjee
Citations: 25
h-index: 1
Agyeya Negi
Agyeya Negi
Citations: 1
h-index: 1

현재의 다중 모드 대규모 언어 모델(MLLM)은 다양한 감각 입력을 처리할 수 있지만, 여전히 특정 모드에 크게 의존하여 추론 능력이 취약하고, 이는 교차 모드 추론 능력 저하로 이어집니다. 본 논문에서는 비디오, 오디오, 이미지 및 텍스트 데이터를 포함하는 3,404개의 샘플로 구성된 C$^3$PO라는 새로운 벤치마크를 소개합니다. 이 벤치마크는 정보 통합(분산된 증거의 결합)과 반사실적 충돌 해결 능력이라는 두 가지 측면을 평가합니다. C$^3$PO는 파트너십 구조와 4단계 설계 방식을 통해 교차 모드 추론 실패가 발생하는 시점과 이유를 정확하게 진단할 수 있도록 합니다. C$^3$PO는 25개의 논리적으로 타당한 템플릿을 사용하여 완전히 자동화된 방식으로 구축되었으며, 이를 통해 인간은 88.64%의 정확도를 달성하는 반면, 최고의 모델(Gemini-3.1-Pro)은 73.17%의 정확도에 그치며, 오픈 소스 모델은 충돌 상황에서 성능이 급격히 저하되는 것을 확인했습니다. 어텐션 분석 결과, 실패 원인의 86~95%가 특정 모드에 대한 의존성에서 비롯되며, 모델은 하나의 모드에 집중하면서 상반되는 증거를 무시하고, 텍스트에 87~95%의 어텐션을 집중하는 것으로 나타났습니다. 중간 레이어의 어텐션 엔트로피는 정확성을 유지하는 탐색이 성공하는 반면, 조기에 붕괴되는 경우에는 실패한다는 것을 예측합니다. 동일한 복잡성의 템플릿 간에 56점의 정확도 차이가 발생하는 것은 성능이 모드의 구조적 역할, 즉 충돌 해결에서의 역할에 따라 달라지며, 단순히 모드 조합에 의해 결정되지 않는다는 점을 보여줍니다. 이러한 결과는 다중 모드 인식이 반드시 강력한 추론 능력을 보장하는 것이 아니며, 아키텍처는 조기 붕괴를 방지하기 위해 지속적인 교차 모드 어텐션을 가능하게 해야 함을 시사합니다.

Original Abstract

Current Multimodal Large Language Models (MLLMs) can process diverse sensory inputs, yet their reasoning remains heavily biased toward a dominant modality, resulting in brittle cross-modal reasoning. We introduce C$^3$PO, a benchmark of 3,404 samples spanning video, audio, image, and text, evaluating two abilities: information composition (fusing dispersed evidence) and counterfactual conflict (resolving deliberate contradictions). C$^3$PO's paired IC/CC structure and four-tier design enable targeted diagnosis of when and why cross-modal reasoning fails. Built through a fully automatic pipeline using 25 logically grounded templates, C$^3$PO reveals that while humans achieve 88.64% accuracy, the best model (Gemini-3.1-Pro) reaches only 73.17%, with open-source models collapsing under conflict. Through attention probes, we find 86-95% of failures stem from modality dominance: models commit to one modality while ignoring contradictory evidence, concentrating 87-95% of attention on text. Mid-layer attention entropy predicts correctness-sustained exploration succeeds, premature collapse fails. The 56-point accuracy gap between equally complex templates reveals that performance depends on modalities' structural roles in conflict resolution, not combinations. These findings show multimodal perception does not guarantee robust reasoning; architectures must enable sustained cross-modal attention to avoid premature

0 Citations
0 Influential
23 Altmetric
115.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!