2608.03733v1 Aug 04, 2026 cs.AI

실패 정보 기반 이미지 자기 증강: 다중 모드 대규모 언어 모델의 자체 개선

Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement

Senkang Hu
Senkang Hu
Citations: 499
h-index: 12
Chi-Min Chan
Chi-Min Chan
Citations: 256
h-index: 7
Yuzhi Zhao
Yuzhi Zhao
Citations: 95
h-index: 6
Wenao Ma
Wenao Ma
Citations: 22
h-index: 3
Wei Xue
Wei Xue
Citations: 55
h-index: 5
Yi-Ting Guo
Yi-Ting Guo
Citations: 2,143
h-index: 26
Sitong Cheng
Sitong Cheng
Citations: 195
h-index: 3
Chunyang Jiang
Chunyang Jiang
Citations: 25
h-index: 4
Zhijian Hou
Zhijian Hou
Citations: 13
h-index: 2
Pingping Zhang
Pingping Zhang
Citations: 15
h-index: 2
Mengyang Wu
Mengyang Wu
Citations: 82
h-index: 6
Yiyang Cai
Yiyang Cai
Citations: 8
h-index: 2

다중 모드 대규모 언어 모델(MLLM)은 시각-언어 작업에서 뛰어난 성능을 보여주었지만, 이러한 발전은 비용이 많이 드는 고품질 다중 모드 데이터에 크게 의존합니다. 자기 증강은 모델이 외부 감독 없이 자체 훈련 데이터를 확장할 수 있는 유망한 대안을 제공합니다. 그러나 기존의 MLLM 자기 증강 방법은 주로 텍스트 중심적이며, 이미지 증강은 상대적으로 연구가 부족하고, 일반적으로 모델의 실제 한계와 일관성이 낮은 일반적이거나 수작업으로 만들어진 변환에 의존합니다. 본 논문에서는 모델 자체의 실패 사례로부터 증강된 이미지를 생성하는 MLLM 자체 개선 프레임워크인 "실패 정보 기반 이미지 자기 증강(Failure-informed Image Self-Augmentation, FISA)"를 제안합니다. 우리의 방법은 시각적으로 어렵지만 답변을 유지하는 이미지 변형을 생성하고, 자체 검사를 통해 유용성을 확인하며, 의미 왜곡을 방지하기 위해 이중 충실도 필터링을 적용합니다. 시각 질의 응답 벤치마크에서의 실험 결과, 제안된 방법은 분포 내 및 분포 외 환경 모두에서 성능을 꾸준히 향상시킵니다. 추가적인 실험을 통해 FISA가 기존의 텍스트 자기 증강 접근 방식과 호환되며, 생성된 샘플이 일반적인 이미지 증강 기준보다 데이터 효율성이 뛰어나고, 제안된 필터링 전략이 실제적으로 효과적임을 확인했습니다.

Original Abstract

Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heavily on large-scale, high-quality multimodal data that are costly to annotate. Self-augmentation offers a promising alternative by enabling models to expand their own training data without external supervision. However, existing MLLM self-augmentation methods are largely text-centric, while image augmentation remains underexplored and typically relies on generic or handcrafted transformations that are weakly aligned with the model's actual incapability. We propose Failure-informed Image Self-Augmentation (\textbf{FISA}), a framework for MLLM self-improvement that constructs augmented images from the model's own failure cases. Our method generates visually challenging yet answer-preserving image complications, verifies their utility through self-examination, and applies dual fidelity filtering to avoid semantic distortion. Experiments on visual question answering benchmarks show that the proposed method consistently improves performance across both in-distribution and out-of-distribution settings. Further experiments validate the compatibility of FISA with existing textual self-augmentation approaches, the superior data efficiency of the synthesized samples over generic image augmentation baselines, and the practical effectiveness of the proposed filtering strategy.

0 Citations
0 Influential
13 Altmetric
65.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!