2607.15216v1 Jul 16, 2026 cs.CV

Symbal: 모델 생성 이미지 설명의 체계적인 불일치 감지

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

M. Varma
M. Varma
Citations: 856
h-index: 11
C. Langlotz
C. Langlotz
Citations: 807
h-index: 14
Jean-Benoit Delbrouck
Jean-Benoit Delbrouck
Citations: 2,821
h-index: 22
S. Ostmeier
S. Ostmeier
Citations: 336
h-index: 7
Akshay S. Chaudhari
Akshay S. Chaudhari
Citations: 365
h-index: 7

다중 모드 대규모 언어 모델(MLLM)은 종종 이미지 설명을 생성할 때 오류를 발생시키며, 이는 이미지와 텍스트 간의 불일치를 초래합니다. 본 연구에서는 MLLM이 생성한 설명에서 발생하는 특정 시각적 특징과 밀접하게 관련된 반복적인 오류인 '체계적인 불일치'에 주목합니다. 본 연구의 목표는 MLLM이 생성한 설명이 포함된 비전-언어 데이터셋에서 이러한 오류를 감지하는 것입니다. 첫 번째 주요 기여로, 당사는 기존 모델을 활용하여 체계적인 불일치를 식별하고 결과를 자연어로 요약하는 구조화되고 이원 단계 방식으로 구성된 'Symbal'을 제안합니다. 두 번째 주요 기여로, 제안된 작업에 대한 자동화 방법의 평가를 위한 벤치마크인 'SymbalBench'를 소개합니다. SymbalBench는 2개의 도메인(자연 이미지 및 의료 이미지)에서 수집된 170만 개의 이미지-텍스트 쌍으로 구성되어 있으며, 주석이 달린 체계적인 불일치를 포함하는 420개의 비전-언어 데이터셋으로 구성됩니다. Symbal은 이 벤치마크에서 뛰어난 성능을 보이며, 전체 데이터셋 중 63.8%에서 체계적인 불일치를 정확하게 식별하여 가장 가까운 기준 모델보다 거의 4배 향상된 성능을 보여줍니다. 또한 SymbalBench에서의 평가 외에도 실제 환경에서의 평가를 통해 (1) Symbal이 4개의 MLLM에서 생성한 설명에서 체계적인 불일치를 정확하게 찾아낼 수 있으며, (2) Symbal은 기존 이미지-텍스트 데이터셋의 품질을 검증하는 강력한 도구임을 확인했습니다. 궁극적으로, 본 연구에서 제시하는 새로운 작업, 방법 및 벤치마크는 사용자가 MLLM이 생성한 설명을 감사하고 중요한 오류를 식별하는 데 도움을 줄 수 있으며, 이를 위해 모델 자체에 대한 접근 권한이 필요하지 않습니다. 관련 코드는 https://github.com/Stanford-AIMI/Symbal 에서 확인할 수 있습니다.

Original Abstract

Multimodal large language models (MLLMs) often introduce errors when generating image captions, resulting in misaligned image-text pairs. Our work focuses on a class of captioning errors that we refer to as systematic misalignments, where a recurring error in MLLM-generated captions is closely associated with the presence of a specific visual feature in the paired image. Given a vision-language dataset with MLLM-generated captions, our aim in this work is to detect such errors, a task we refer to as systematic misalignment detection. As our first key contribution, we present Symbal, which utilizes a structured, dual-stage setup with off-the-shelf foundation models to identify systematic misalignments and summarize results in natural language. As our second key contribution, we introduce SymbalBench, a benchmark designed to evaluate automated methods on our proposed task. SymbalBench consists of 1.7 million image-text pairs from two domains (natural and medical images), organized into 420 vision-language datasets with annotated systematic misalignments. Symbal exhibits strong performance on this benchmark, correctly identifying systematic misalignments in 63.8% of datasets, a nearly 4x improvement over the closest baseline. We supplement our evaluations on SymbalBench with real-world evaluations, showing that (1) Symbal can accurately surface systematic misalignments in captions generated by four MLLMs and (2) Symbal is a powerful tool for auditing off-the-shelf image-caption datasets. Ultimately, our novel task, method, and benchmark can aid users with auditing MLLM-generated captions and identifying critical errors, without requiring access to the underlying MLLM. Code is available at https://github.com/Stanford-AIMI/Symbal.

0 Citations
0 Influential
31 Altmetric
0.0 Score
Original PDF
0

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!