2608.04111v1 Aug 04, 2026 cs.CV

GEB-벤치: 다양한 관점에서 설명되는 추상적 구조

GEB-Bench: Abstract Structures Told in Many Voices

Yun Peng
Yun Peng
Citations: 78
h-index: 3
Tao Xie
Tao Xie
Citations: 196
h-index: 5
Tong Zhang
Tong Zhang
Citations: 0
h-index: 0
Zhiyuan Shi
Zhiyuan Shi
Citations: 0
h-index: 0

모델이 강과 번개를 보고 이들이 동일한 구조를 가지고 있다는 것을 인식할 수 있을까요? 본 논문에서는 괴델, 에셔, 바흐의 정신을 담은 벤치마크인 GEB-벤치를 소개합니다. GEB-벤치의 기본 단위는 자기 참조, 이상 루프, 모비우스 트위스트와 같은 추상적인 구조적 모티프입니다. 각 모티프는 여러 관점에서 설명됩니다: 구조를 구성하는 자연 현상, 기계적으로 검증 가능한 장치를 통해 이를 구현하는 민담, 수학적 정리, 그리고 프로그램 코드로 표현된 골격; 표면 파라미터는 간섭 변수로 간주되어 점수에 반영되지 않습니다. 모티프, 관점, 그리고 이들 사이의 구조적 변화는 작은 다중 모달 범주를 형성하며, GEB-벤치의 작업은 이러한 범주에 대한 질문입니다. 12개의 공개 및 독점 모델을 평가한 결과, 추상화 실패가 일관된 패턴을 보임을 확인했습니다. 핵심적인 발견은 인식과 관점 간 매핑의 격차입니다: 모델은 특정 관점에서 구조를 더 잘 식별하지만, 이를 다른 관점으로 확장하는 데는 어려움을 겪습니다. 모든 모델이 이러한 제약을 받으며, 이 격차를 줄일 수 있는 강력한 매핑은 최첨단 모델에서만 나타납니다. 두 가지 패턴이 이를 뒷받침합니다. 오류는 측정된 인지적 구조보다 설계된 형식 기하학에 더 밀접하게 연관되어 있으며, 서로 다른 벤더의 최첨단 모델들은 동일한 잘못된 답변으로 수렴하는 경향이 있습니다. 또한, 표면 복잡성은 구조를 인식하는 모든 모델에 부담을 주며, 용량은 면역력을 제공하기보다는 여유 공간을 확보하는 데 사용됩니다. GEB-벤치는 완전히 생성 가능하며, 파이프라인과 함께 공개되었습니다.

Original Abstract

Can a model look at a river delta and a lightning bolt and see that they share a structure? We introduce GEB-Bench, a benchmark whose unit is an abstract structural motif--self-reference, a strange loop, a Mobius twist--in the spirit of Godel, Escher, Bach. Each motif is told in several voices: a natural scene whose composition is the structure, a folk story whose telling enacts it through a mechanically checkable form device, a mathematical theorem, and a programmatic skeleton; surface parameters are declared nuisance variables and never scored. Motifs, voices, and the structural changes between them form a small cross-modal category, and GEB-Bench's tasks are its questions. Evaluating twelve open and proprietary models, we find that abstraction failure is lawful. The central finding is a gap between recognition and cross-voice mapping: models identify a structure within one voice far better than they carry it across voices; every model pays this tax, and mapping strong enough to narrow it appears only at the frontier tier. Two patterns support it. Errors align more strongly with the designed formal geometry than with measured perceptual geometries, and frontier models from different vendors converge on the same wrong answers; and surface complexity taxes every model that reads structure, with capacity buying headroom rather than immunity. GEB-Bench is fully generative and released with its pipeline.

0 Citations
0 Influential
2.5 Altmetric
12.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!