2601.11675v3 Jan 16, 2026 cs.CV

인간의 장면 이해 메타머 생성

Generating metamers of human scene understanding

Ritik Raina
Ritik Raina
Citations: 0
h-index: 0
Abe Leite
Abe Leite
Citations: 25
h-index: 1
Alexandros Graikos
Alexandros Graikos
Citations: 614
h-index: 10
Seoyoung Ahn
Seoyoung Ahn
Citations: 601
h-index: 13
Dimitris Samaras
Dimitris Samaras
Citations: 71
h-index: 4
G. Zelinsky
G. Zelinsky
Citations: 6,612
h-index: 43

인간의 시각은 시각 주변부에서 얻은 저해상도 '전반적인 정보'와 고정된 위치에서 얻은 희소하지만 고해상도 정보를 결합하여 시각 장면을 일관성 있게 이해합니다. 본 논문에서는 인간의 잠재적인 장면 표현과 일치하는 장면을 생성하는 도구인 MetamerGen을 소개합니다. MetamerGen은 주변부에서 얻은 장면의 전반적인 정보와 장면 시청 시 고정점에서 얻은 정보를 결합하여, 사람이 장면을 시청한 후 이해하는 내용과 유사한 이미지 메타머를 생성하는 잠재 확산 모델입니다. 고해상도와 저해상도(즉, '망막 중심부') 입력을 모두 사용하여 이미지를 생성하는 것은 새로운 이미지-이미지 합성 문제입니다. 우리는 이 문제를 해결하기 위해, 망막 중심부 장면의 상세한 특징과 장면의 맥락을 나타내는 주변부에서 손실된 특징을 결합하는 DINOv2 토큰으로 구성된 이중 스트림 표현을 도입했습니다. MetamerGen이 생성한 이미지와 인간의 잠재적인 장면 표현 간의 지각적 일관성을 평가하기 위해, 참가자들에게 생성된 이미지와 원본 이미지 간에 '동일' 또는 '다름'을 판단하도록 하는 행동 실험을 수행했습니다. 이를 통해, 참가자들이 인식하는 장면 표현에 대한 메타머인 장면 생성을 식별했습니다. MetamerGen은 장면 이해를 연구하는 데 유용한 도구입니다. 우리의 개념 증명 분석 결과, 인간의 판단에 기여하는 시각 처리의 여러 단계에서 특정 특징들이 작용하는 것을 확인했습니다. MetamerGen은 임의의 고정 위치에 조건화되어도 메타머를 생성할 수 있지만, 생성된 장면이 시청자의 고정 영역에 조건화되었을 때 고수준의 의미적 일관성이 메타머 현상을 가장 강력하게 예측한다는 것을 발견했습니다.

Original Abstract

Human vision combines low-resolution "gist" information from the visual periphery with sparse but high-resolution information from fixated locations to construct a coherent understanding of a visual scene. In this paper, we introduce MetamerGen, a tool for generating scenes that are aligned with latent human scene representations. MetamerGen is a latent diffusion model that combines peripherally obtained scene gist information with information obtained from scene-viewing fixations to generate image metamers for what humans understand after viewing a scene. Generating images from both high and low resolution (i.e. "foveated") inputs constitutes a novel image-to-image synthesis problem, which we tackle by introducing a dual-stream representation of the foveated scenes consisting of DINOv2 tokens that fuse detailed features from fixated areas with peripherally degraded features capturing scene context. To evaluate the perceptual alignment of MetamerGen generated images to latent human scene representations, we conducted a same-different behavioral experiment where participants were asked for a "same" or "different" response between the generated and the original image. With that, we identify scene generations that are indeed metamers for the latent scene representations formed by the viewers. MetamerGen is a powerful tool for understanding scene understanding. Our proof-of-concept analyses uncovered specific features at multiple levels of visual processing that contributed to human judgments. While it can generate metamers even conditioned on random fixations, we find that high-level semantic alignment most strongly predicts metamerism when the generated scenes are conditioned on viewers' own fixated regions.

0 Citations
0 Influential
21.5 Altmetric
107.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!