2607.27670v1 Jul 30, 2026 cs.CV

JigShape: 지그소 퍼즐을 활용한 VLMs의 시각-기하 추론 능력 평가

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Shawn Li
Shawn Li
Citations: 4
h-index: 1
Jiate Li
Jiate Li
Citations: 15
h-index: 3
Youxuan Qin
Youxuan Qin
Citations: 124
h-index: 4
Roger Zimmermann
Roger Zimmermann
Citations: 50
h-index: 5
Vicente Ordonez
Vicente Ordonez
Citations: 57
h-index: 4
Wei Yang
Wei Yang
Citations: 19
h-index: 2
Jike Zhong
Jike Zhong
Citations: 251
h-index: 5
Jiawei Yang
Jiawei Yang
Citations: 1,065
h-index: 9
Ryan A. Rossi
Ryan A. Rossi
Citations: 46
h-index: 2
Franck Dernoncourt
Franck Dernoncourt
Citations: 49
h-index: 4
Zhengzhong Tu
Zhengzhong Tu
Citations: 205
h-index: 9
Mohit Bansal
Mohit Bansal
Citations: 231
h-index: 7
Yue Zhao
Yue Zhao
Citations: 115
h-index: 5

지그소 퍼즐 풀이는 시각적 내용과 기하학적 제약을 동시에 고려해야 하지만, 기존 벤치마크는 반복적인 질감이 있는 영역에서 모호한 정답을 생성하는 직사각형 조각을 사용합니다. 본 논문에서는 extit{ ext{JigShape}}, 즉 탭-앤-블랭크 연결 방식으로 제작된 조각을 사용하는 벤치마크를 소개합니다. 이 방식은 기하학적 제약이 강한 국소적인 호환성을 제공하며, 시각적 내용과 결합하여 명확한 정답을 얻을 수 있습니다. 95,000개의 인스턴스와 네 가지 그리드 밀도(4x4부터 16x16까지)를 사용하여 분석한 결과, extbf{제로샷 VLMs는 기하학적 추론 능력이 부족}합니다. 다섯 개의 최첨단 모델 중 GPT-5.5만이 4x4 퍼즐에서 무작위 기준선을 넘어서는 성능을 보였으며, 나머지 모델은 우연 수준의 성능에 머물렀습니다. 지도 학습 기반 미세 조정(supervised fine-tuning)을 통해 4x4 퍼즐에서는 97% 이상의 정확도를 달성했지만, extbf{모든 모델이 더 큰 그리드 크기에서 성능 저하}를 보였습니다. GPT-5.5는 8x8 퍼즐에서 70%에서 거의 무작위 수준으로 떨어졌으며, 미세 조정된 모델조차도 12x16 퍼즐에서 5% 이하의 성능을 보였습니다. 이러한 '확장성 제한(scaling cliff)' 현상은 현재 아키텍처가 조각 수가 증가함에 따라 일관된 제약 조건 만족을 유지할 수 없음을 시사합니다. ext{JigShape}는 컴퓨터 비전-언어 모델에게 확장 가능한 기하학적 추론 능력을 향상시키는 과제로 제시합니다.

Original Abstract

Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that create ambiguous ground truth in texture-repeated regions. We introduce \textit{\ours{}}, a benchmark with tab-and-blank interlocking pieces where geometric constraints provide strong local compatibility requirements that, combined with visual content, yield unambiguous ground truth. Across 95K instances at four grid densities (4$\times$4 to 16$\times$16), we find that \textbf{zero-shot VLMs largely lack geometric reasoning}: only one of five frontier models (GPT-5.5) exceeds random baseline on 4$\times$4 puzzles, while all others perform at chance level. While supervised fine-tuning achieves $>$97\% on 4$\times$4, \textbf{all models collapse on larger grids}: GPT-5.5 drops from 70\% to near-random on 8$\times$8, and even fine-tuned models fall below 5\% on 12$\times$12. This ``scaling cliff'' suggests current architectures cannot maintain consistent constraint satisfaction as the number of pieces increases. \ours{} establishes scalable geometric reasoning as an open challenge for vision-language models.

0 Citations
0 Influential
4.5 Altmetric
22.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!