2607.25294v1 Jul 28, 2026 cs.CV

CLBench-V: 정량화된 지식을 활용한 다중 모드 컨텍스트 학습 평가

CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

Lai Wei
Lai Wei
Shanghai Jiao Tong University
Citations: 227
h-index: 7
Weiran Huang
Weiran Huang
Citations: 348
h-index: 10
Yue Wang
Yue Wang
Citations: 108
h-index: 5
Jiapeng Li
Jiapeng Li
Citations: 15
h-index: 2
Rui Hu
Rui Hu
Citations: 40
h-index: 2
Chengqi Li
Chengqi Li
Citations: 0
h-index: 0

실제 환경에서의 작업은 모델이 사전 훈련된 지식에만 의존하는 것이 아니라, 작업별 맥락에서 학습하도록 요구하는 경우가 많습니다. 최근 연구에서는 이러한 능력을 '컨텍스트 학습(context learning)'이라고 부르며 강조하고 있지만, 기존의 평가 방법은 주로 텍스트 기반 컨텍스트에 초점을 맞추고 있습니다. 하지만 많은 실제 환경에서 모델이 학습해야 할 맥락은 다중 모드 형태를 취합니다. 예를 들어, 과학적 발견은 그림과 표를 통해 전달되고, 금융 지표는 다양한 보고서에 분산되어 있으며, 공간적인 의사 결정은 지도, 이미지 또는 웹 페이지에 따라 달라집니다. 본 연구에서는 이러한 문제를 해결하기 위해 세 가지 측면(맥락 정량화, 새로운 정보 적용, 새로운 지식 습득)을 중심으로 다중 모드 컨텍스트 학습을 평가하는 벤치마크인 CLBench-V를 소개합니다. CLBench-V는 기존의 공개 벤치마크와 함께 과학, 금융, 긴 문서 이해, 공간 추론 및 웹 기반 시각적 질의응답 등 다양한 분야를 아우르는 새로운 데이터셋을 결합합니다. 또한, 도메인별 컨텍스트 학습 작업을 구축하는 데 드는 비용을 줄이기 위해 자동으로 생성하고 필터링하는 절차를 사용하여 새로 구축된 데이터셋을 구성했습니다. 3,443개의 인스턴스와 최근의 다중 모드 모델 6개를 대상으로 실험한 결과, 최고 점수는 0.2847에 불과하여 다중 모드 컨텍스트 학습은 아직 발전 가능성이 높다는 것을 보여줍니다. 또한, InternVL3.5-30B-A3B는 맥락 정량화 및 새로운 지식 습득에서 가장 좋은 성능을 보였으며, Qwen3.5-Plus는 새로운 정보 적용에서 가장 뛰어난 성능을 보였습니다. 본 연구에서는 평가자의 신뢰성, 컨텍스트 길이, 이미지 개수 및 대표적인 실패 사례에 대한 추가 분석을 수행했습니다. 코드 및 관련 자료는 https://github.com/IamLihua/CLBench-V 에서 확인할 수 있습니다.

Original Abstract

Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus on textual contexts. In many practical settings, however, the context to be learned from is multimodal: scientific findings are conveyed through figures and tables, financial indicators are scattered across converted reports, and spatial decisions depend on maps, scenes, or web pages. We introduce CLBench-V, a benchmark for multimodal context learning that addresses the difficulty of localizing where context use breaks down by organizing tasks around three dimensions: context grounding, new information application, and new knowledge learning. CLBench-V combines converted public benchmarks with newly constructed datasets spanning domains such as science, finance, long-document understanding, spatial reasoning, and web-based visual question answering. To reduce the cost of constructing domain-specific context-learning tasks, we further use automated construction and filtering procedures for our newly built datasets. Across 3,443 instances and six recent multimodal models, the best overall score is only 0.2847, indicating that multimodal context learning remains far from saturated. Moreover, InternVL3.5-30B-A3B performs best on context grounding and new knowledge learning, while Qwen3.5-Plus performs best on new information application. We further analyze judge reliability, context length, image count, and representative failure cases. Code is available at https://github.com/IamLihua/CLBench-V.

0 Citations
0 Influential
21 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!