2603.23867v1 Mar 25, 2026 cs.LG

대규모 시각-언어 모델(VLM)은 얼마나 안정적으로 추론할 수 있는가? 신경-기호 접근 방식 연구

Can VLMs Reason Robustly? A Neuro-Symbolic Investigation

Antonio Vergari
Antonio Vergari
University of Edinburgh
Citations: 2,412
h-index: 25
Weixin Chen
Weixin Chen
Citations: 11
h-index: 2
Han Zhao
Han Zhao
Citations: 45
h-index: 2

시각-언어 모델(VLM)은 다양한 추론 작업에 적용되어 왔지만, 데이터 분포 변화 하에서 VLM이 얼마나 안정적으로 추론할 수 있는지는 여전히 불분명합니다. 본 연구에서는, 시각적 입력 분포는 변화하지만, 근본적인 예측 규칙은 동일하게 유지되는 경우의 데이터 분포 변화를 연구합니다. 이 질문을 탐구하기 위해, 우리는 모델이 이미지와 이미지 내 객체 개념에 대한 논리 규칙을 기반으로 질문에 답해야 하는 시각적 연역 추론 작업을 고려합니다. 실험 결과, 경사 기반의 엔드 투 엔드 학습을 통해 미세 조정된 VLM은 동일한 데이터 분포 내에서는 높은 정확도를 달성하지만, 데이터 분포 변화 하에서는 일반화에 실패하는 것으로 나타났습니다. 이는 미세 조정이 반드시 근본적인 추론 기능을 유도하지는 않는다는 것을 시사합니다. 이러한 점을 고려하여, 우리는 인식과 추론을 분리하는 신경-기호 관점을 제시합니다. 그러나, 최근의 신경-기호 접근 방식 중 추론에 블랙박스 구성 요소를 사용하는 방식은 여전히 작업에 따라 일관성 없는 안정성을 보이는 것을 확인했습니다. 이러한 문제를 해결하기 위해, 우리는 VLM 기반의 개념 인식과 회로 기반의 기호 추론을 결합한 신경-기호 방법인 VLC를 제안합니다. 특히, 작업 규칙은 기호 프로그램, 즉 회로로 컴파일되어 VLM에 의해 인식된 객체 개념에 대해 정확하게 규칙을 실행합니다. 세 가지 서로 다른 규칙 집합을 가진 시각적 연역 추론 작업에 대한 실험 결과, VLC는 데이터 분포 변화 하에서 일관되게 강력한 성능을 보이는 것으로 나타났으며, 이는 VLC가 안정적인 추론을 지원하는 능력을 보여줍니다.

Original Abstract

Vision-Language Models (VLMs) have been applied to a wide range of reasoning tasks, yet it remains unclear whether they can reason robustly under distribution shifts. In this paper, we study covariate shifts in which the perceptual input distribution changes while the underlying prediction rules do not. To investigate this question, we consider visual deductive reasoning tasks, where a model is required to answer a query given an image and logical rules defined over the object concepts in the image. Empirically, we find that VLMs fine-tuned through gradient-based end-to-end training can achieve high in-distribution accuracy but fail to generalize under such shifts, suggesting that fine-tuning does not reliably induce the underlying reasoning function. This motivates a neuro-symbolic perspective that decouples perception from reasoning. However, we further observe that recent neuro-symbolic approaches that rely on black-box components for reasoning can still exhibit inconsistent robustness across tasks. To address this issue, we propose VLC, a neuro-symbolic method that combines VLM-based concept recognition with circuit-based symbolic reasoning. In particular, task rules are compiled into a symbolic program, specifically a circuit, which executes the rules exactly over the object concepts recognized by the VLM. Experiments on three visual deductive reasoning tasks with distinct rule sets show that VLC consistently achieves strong performance under covariate shifts, highlighting its ability to support robust reasoning.

1 Citations
0 Influential
12.5 Altmetric
63.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!