2608.06142v1 Aug 06, 2026 cs.CV

작품 및 사진의 구성 분석을 위한 시각적 표현 학습

Learning visual representations for compositional analysis of artworks and photographs

T. Tuytelaars
T. Tuytelaars
Citations: 57,177
h-index: 76
F. Behrad
F. Behrad
Citations: 18
h-index: 3
Johan Wagemans
Johan Wagemans
Citations: 602
h-index: 10

구성, 즉 시각적 요소들의 의도적인 배치 방식은 예술 작품에서 의미, 감정, 그리고 미적 품질이 전달되는 데 핵심적인 역할을 하지만, 시각적 이해의 여러 측면 중 가장 체계화되지 않은 영역 중 하나입니다. 기존 연구에서는 의미 있는 구성 표현 학습에 있어 지속적인 격차가 존재하며, 이는 의미론적 편향 때문이며, 인간에게 영감을 받은 접근 방식이 중요한 해결책이 될 수 있다고 지적합니다. 본 연구에서는 구성 분석을 위한 두 가지 병렬적인 방법을 비교합니다. 첫 번째는 인지적 그룹핑에 기반한 인간에게 영감을 받은 방법이고, 두 번째는 최근 대규모 구성 데이터셋을 활용하여 미세 조정된 기본 모델입니다. 인간에게 영감을 받은 접근 방식은 객체 중심 모델을 사용하여 영역 수준의 분해를 수행하고, 그래프 어텐션 네트워크를 사용하여 요소 간의 공간적 관계를 파악합니다. 두 가지 방법 모두 구성 점수/범주 예측, 구성 이미지 검색 및 시각적 주의 집중 감지를 위해 평가됩니다. 인코더를 고정했을 때, 인간에게 영감을 받은 방법은 경쟁력 있는 성능을 보이면서도 해석 가능성을 유지합니다. 충분한 데이터가 확보되어 미세 조정을 수행할 수 있을 경우, 대규모 자기 지도 학습 모델은 훨씬 더 뛰어난 성능을 보이지만, 해석 가능성과 다양한 영역으로의 일반화 능력 측면에서는 단점을 가집니다.

Original Abstract

Composition, the deliberate arrangement of visual elements, is central to how meaning, emotion, and aesthetic quality are conveyed in artwork, yet it remains among the least formalized dimensions of visual understanding. Prior work highlights a persistent gap in learning meaningful compositional representations, attributing it to semantic bias and suggesting that human-inspired approaches may be key. We compare two parallel paradigms for composition analysis: a human-inspired method grounded in perceptual grouping, and fine-tuned foundation models enabled by recent large-scale compositional datasets. The human-inspired approach uses object-centric models for region-level decomposition and a graph attention network to capture spatial relationships between elements. Both paradigms are evaluated on composition score/category prediction, compositional image retrieval, and visual saliency detection. With frozen encoders, the human-inspired method achieves competitive performance while remaining interpretable. When sufficient data enables fine-tuning, large self-supervised models outperform significantly, but at the cost of interpretability and cross-domain generalization.

0 Citations
0 Influential
30 Altmetric
150.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!