2607.27747v1 Jul 30, 2026 cs.CL

LVLM은 시각적 착시의 진실을 밝힐 수 있는가? 인지 및 추론 능력 분석

Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities

Yulan Hu
Yulan Hu
Citations: 74
h-index: 5
Rong Yin
Rong Yin
Citations: 1
h-index: 1
Liang Zhao
Liang Zhao
Citations: 0
h-index: 0
Jiaqing Lyu
Jiaqing Lyu
Citations: 0
h-index: 0
Kexin Tang
Kexin Tang
Citations: 97
h-index: 4
Zecheng Fang
Zecheng Fang
Citations: 73
h-index: 5
Da Li
Da Li
Citations: 21
h-index: 2
Jianing Li
Jianing Li
Citations: 0
h-index: 0

대규모 비전-언어 모델(LVLM)은 추론 능력을 통합하여 인지 성능을 새로운 수준으로 끌어올렸습니다. 그러나 기존 평가 방식은 주로 시각적 인식에만 초점을 맞추거나 수학 또는 코딩과 같은 특정 영역에 의존합니다. 개방형 환경과 일치하는 추론 능력에 대한 평가는 여전히 필요하며, 특히 시각적 인식과 추론을 함께 고려하는 평가가 중요합니다. 이러한 격차를 해소하기 위해, 우리는 시각적 착시를 진단 도구로 활용하여 LVLM을 평가하는 방법을 제안합니다. 시각적 착시는 인간의 시각 시스템이 객관적인 신호를 잘못 해석하여 현실과 다른 이해를 초래하는 현상입니다. 우리는 실제 환경에서 수집된 다양한 주석 처리된 질문-답변 쌍을 포함하는 시각적 착시 이미지 벤치마크인 IllusionReasoning을 구축했습니다. IllusionReasoning을 기반으로, 광범위한 LVLM의 추론 능력이 주장만큼 발전하지 않았음을 보여줍니다. 본 연구는 LVLM에 대한 새로운 통찰력을 제공하며, 향후 최적화를 위한 방향을 제시합니다.

Original Abstract

Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evaluations either focus solely on perception or rely on specific domains such as maths or coding. Evaluation for reasoning capabilities that align with an open-world environment is still required, especially one that considers perception and reasoning jointly. To bridge this gap, we propose to evaluate LVLMs by exploiting visual illusions as a diagnostic tool. Visual illusions are phenomena in which the human visual system misinterprets objective signals, resulting in an understanding that deviates from reality. We constructed IllusionReasoning, a benchmark of illusion images collected from the real world, incorporating diverse annotated question-answer pairs. Based on IllusionReasoning, we show that the reasoning capabilities of a wide range of LVLMs are not as advanced as claimed. Our work provides new insights into LVLMs and offers future direction for optimisation.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!