2602.03038v1 Feb 03, 2026 cs.CV

인식과 추론의 경계에서: 프로그램인가, 언어인가? - 봉가드 문제 연구

Bongards at the Boundary of Perception and Reasoning: Programs or Language?

Cassidy Langenfeld
Cassidy Langenfeld
Citations: 0
h-index: 0
Claas Beger
Claas Beger
Citations: 23
h-index: 2
Gloria Geng
Gloria Geng
Citations: 2
h-index: 1
Wasu Top Piriyakulkij
Wasu Top Piriyakulkij
Cornell University
Citations: 74
h-index: 3
Keya Hu
Keya Hu
Citations: 117
h-index: 3
Yewen Pu
Yewen Pu
Citations: 244
h-index: 6
Kevin Ellis
Kevin Ellis
Citations: 0
h-index: 0

비전-언어 모델(VLMs)은 자연 이미지에 대한 설명 생성 또는 이미지에 대한 상식적인 질문에 답변하는 등 일상적인 시각 작업에서 상당한 발전을 이루었습니다. 하지만 인간은 놀라운 능력을 가지고 있으며, 이는 봉가드 문제로 알려진 고전적인 시각 추론 과제 세트를 통해 엄격하게 테스트됩니다. 본 연구에서는 이러한 문제를 해결하기 위한 신경-기호 접근 방식을 제시합니다. 봉가드 문제에 대한 가설적인 해결 규칙이 주어지면, LLM을 활용하여 해당 규칙에 대한 매개변수화된 프로그램 표현을 생성하고, 베이즈 최적화를 사용하여 매개변수 조정을 수행합니다. 저희는 제시된 방법을, 정답 규칙이 주어졌을 때 봉가드 문제 이미지 분류에 적용하고, 또한 문제를 처음부터 해결하는 데 적용하여 평가했습니다.

Original Abstract

Vision-Language Models (VLMs) have made great strides in everyday visual tasks, such as captioning a natural image, or answering commonsense questions about such images. But humans possess the puzzling ability to deploy their visual reasoning abilities in radically new situations, a skill rigorously tested by the classic set of visual reasoning challenges known as the Bongard problems. We present a neurosymbolic approach to solving these problems: given a hypothesized solution rule for a Bongard problem, we leverage LLMs to generate parameterized programmatic representations for the rule and perform parameter fitting using Bayesian optimization. We evaluate our method on classifying Bongard problem images given the ground truth rule, as well as on solving the problems from scratch.

0 Citations
0 Influential
3 Altmetric
15.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!