대규모 시각-언어 모델(VLM)은 얼마나 안정적으로 추론할 수 있는가? 신경-기호 접근 방식 연구
Can VLMs Reason Robustly? A Neuro-Symbolic Investigation
시각-언어 모델(VLM)은 다양한 추론 작업에 적용되어 왔지만, 데이터 분포 변화 하에서 VLM이 얼마나 안정적으로 추론할 수 있는지는 여전히 불분명합니다. 본 연구에서는, 시각적 입력 분포는 변화하지만, 근본적인 예측 규칙은 동일하게 유지되는 경우의 데이터 분포 변화를 연구합니다. 이 질문을 탐구하기 위해, 우리는 모델이 이미지와 이미지 내 객체 개념에 대한 논리 규칙을 기반으로 질문에 답해야 하는 시각적 연역 추론 작업을 고려합니다. 실험 결과, 경사 기반의 엔드 투 엔드 학습을 통해 미세 조정된 VLM은 동일한 데이터 분포 내에서는 높은 정확도를 달성하지만, 데이터 분포 변화 하에서는 일반화에 실패하는 것으로 나타났습니다. 이는 미세 조정이 반드시 근본적인 추론 기능을 유도하지는 않는다는 것을 시사합니다. 이러한 점을 고려하여, 우리는 인식과 추론을 분리하는 신경-기호 관점을 제시합니다. 그러나, 최근의 신경-기호 접근 방식 중 추론에 블랙박스 구성 요소를 사용하는 방식은 여전히 작업에 따라 일관성 없는 안정성을 보이는 것을 확인했습니다. 이러한 문제를 해결하기 위해, 우리는 VLM 기반의 개념 인식과 회로 기반의 기호 추론을 결합한 신경-기호 방법인 VLC를 제안합니다. 특히, 작업 규칙은 기호 프로그램, 즉 회로로 컴파일되어 VLM에 의해 인식된 객체 개념에 대해 정확하게 규칙을 실행합니다. 세 가지 서로 다른 규칙 집합을 가진 시각적 연역 추론 작업에 대한 실험 결과, VLC는 데이터 분포 변화 하에서 일관되게 강력한 성능을 보이는 것으로 나타났으며, 이는 VLC가 안정적인 추론을 지원하는 능력을 보여줍니다.
Vision-Language Models (VLMs) have been applied to a wide range of reasoning tasks, yet it remains unclear whether they can reason robustly under distribution shifts. In this paper, we study covariate shifts in which the perceptual input distribution changes while the underlying prediction rules do not. To investigate this question, we consider visual deductive reasoning tasks, where a model is required to answer a query given an image and logical rules defined over the object concepts in the image. Empirically, we find that VLMs fine-tuned through gradient-based end-to-end training can achieve high in-distribution accuracy but fail to generalize under such shifts, suggesting that fine-tuning does not reliably induce the underlying reasoning function. This motivates a neuro-symbolic perspective that decouples perception from reasoning. However, we further observe that recent neuro-symbolic approaches that rely on black-box components for reasoning can still exhibit inconsistent robustness across tasks. To address this issue, we propose VLC, a neuro-symbolic method that combines VLM-based concept recognition with circuit-based symbolic reasoning. In particular, task rules are compiled into a symbolic program, specifically a circuit, which executes the rules exactly over the object concepts recognized by the VLM. Experiments on three visual deductive reasoning tasks with distinct rule sets show that VLC consistently achieves strong performance under covariate shifts, highlighting its ability to support robust reasoning.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.