2608.06501v1 Aug 06, 2026 cs.AI

MLLM은 창의적 도약을 이해할 수 있는가? 교차 개념 이해를 위한 C4 프레임워크 소개

Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Shi Feng
Shi Feng
Citations: 105
h-index: 5
Xiaocui Yang
Xiaocui Yang
Citations: 814
h-index: 13
Daling Wang
Daling Wang
Citations: 2,694
h-index: 25
Ming Wang
Ming Wang
Northeastern University
Citations: 435
h-index: 7
Tingna Xie
Tingna Xie
Citations: 2
h-index: 1
Yuqi Zhang
Yuqi Zhang
Citations: 0
h-index: 0
Xiangju Li
Xiangju Li
Citations: 377
h-index: 9

MLLM(대규모 언어 모델)의 창의적 능력은 디자인, 커뮤니케이션, 교육 및 인간-AI 협업에 중요하지만, 정확도 중심적인 작업과 비교했을 때 명확한 목표와 보상 신호가 부족하여 평가하기 어렵습니다. 교차 개념 이해는 수동적 창의성의 핵심 인지 능력이며, 이는 정보 수신자가 명시적으로 드러나지 않지만 의미 있는 개념 간의 관계로부터 의도된 의미를 파악할 수 있도록 합니다. 우리는 문제 구성 방식을 교차 개념 인코딩으로, 모델 추론을 교차 개념 디코딩으로 정의합니다. 본 논문에서는 중국어 관용구(체요우) 기반의 교차 개념 창의성을 평가하기 위한 인지 기반 프레임워크인 C4를 소개합니다. C4의 인코딩 구성 요소는 대상 슬롯을 수동으로 주석이 달린, 외부 검토를 거친 교차 개념 네트워크 내에서 이미지화 가능한 대체 개념으로 매핑하여 명확한 구조와 난이도(브릿지 개수 및 깊이로 측정)를 갖춘 일괄 생성과 정확한 정답을 가능하게 합니다. 본 프레임워크를 사용하여 184개의 합성 항목과 온라인 소스에서 수집된 37개의 인간이 만든 교차 개념 체요우 그림으로 구성된 C4 평가 세트(C4-Eval)를 구현했습니다. 우리는 수집된 그림에 대한 교차 개념 관계, 브릿지 경로 및 추론 과정을 수동으로 구성하고 검토합니다. 각 C4-Eval 항목은 5가지 작업 설정에서 구현되어 총 884개의 주요 답변 회수 사례를 생성합니다. 평가된 10개의 MLLM 모델 중 가장 성능이 좋은 폐쇄형 모델의 경우, 주요 정확도가 각각 50.7% 및 48.0%에 달하는 반면, 오픈 소스 모델은 훨씬 낮은 수준을 보입니다. 후보 제약 조건은 정확도를 크게 향상시키지만, 브릿지 힌트와 설명 요청은 미미한 개선 효과만 제공합니다. 이러한 결과는 현재 MLLM이 교차 개념 관계를 통해 창의적으로 인코딩된 의미를 어떻게 디코딩하는지에 대한 상당한 격차가 있음을 보여줍니다. 관련 코드는 추가 자료에 포함되어 있습니다.

Original Abstract

Creative capabilities of MLLMs matter in design, communication, education, and human--AI collaboration, yet remain difficult to evaluate because explicit targets and reward signals are scarce compared with accuracy-oriented tasks. Cross-concept understanding is a core cognitive capacity underlying receptive creativity. It enables a perceiver to recover intended meaning from non-obvious but meaningful conceptual relations. We operationalize item construction as cross-concept encoding and model inference as cross-concept decoding. We introduce C4, a cognition-inspired evaluation framework for Chengyu (Chinese idiom)-based Cross-Concept Creativity. Its encoding component maps target slots to imageable substitute concepts along bridge paths in a manually annotated and third-party-reviewed cross-concept network, enabling batch generation with explicit structure, difficulty indexed by bridge count and depth, and exact answers. Using this framework, we instantiate the C4 Evaluation Set (C4-Eval), comprising 184 synthetic items and 37 human-created cross-concept chengyu figures collected from online sources. We manually construct and review cross-concept relations, bridge paths, and reasoning processes for the collected figures. Each C4-Eval item is instantiated in five task settings, yielding 884 primary answer-recovery cases. Across ten evaluated MLLMs, the strongest closed models reach 50.7% and 48.0% primary accuracy, while open-source models remain substantially lower. Candidate constraints improve accuracy sharply, but bridge hints and explanation requests provide only modest gains. These results expose a substantial gap in how current MLLMs decode creatively encoded meaning through cross-concept relations. The code is in the supplementary material.

0 Citations
0 Influential
12.5 Altmetric
62.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!