2604.02031v1 Apr 02, 2026 cs.CV

희소성 인지 자동 인코딩: 공간적으로 불균형한 데이터 복원

Rare-Aware Autoencoding: Reconstructing Spatially Imbalanced Data

A. Garcia
A. Garcia
Citations: 30
h-index: 2
J. V. Gemert
J. V. Gemert
Citations: 35
h-index: 3
Daan Brinks
Daan Brinks
Citations: 431
h-index: 7
Nergis Tömen
Nergis Tömen
Citations: 71
h-index: 3

자동 인코더는 이미지 콘텐츠의 공간적 불균일성으로 인해 어려움을 겪을 수 있습니다. 이는 의료 영상, 생물학 및 물리학 분야에서 흔히 나타나는 현상으로, 유용한 패턴이 특정 이미지 좌표에 드물게 나타나지만, 대부분의 샘플에서 배경이 지배적이어서 복원 과정이 다수 패턴에 편향될 수 있습니다. 실제로, 자동 인코더는 지배적인 패턴에 편향되어 세부 사항이 손실되고 공간 데이터 불균형이 심한 경우 특히 희귀한 공간 입력에 대해 흐릿한 복원을 초래합니다. 우리는 공간적 불균형 문제를 해결하기 위해 두 가지 상호 보완적인 방법을 사용합니다. (i) 통계적으로 드문 공간 위치에 더 높은 가중치를 부여하는 자체 엔트로피 기반 손실 함수와 (ii) Sample Propagation이라는 리플레이 메커니즘을 사용하여 학습 중에 배치 내에서 재구성하기 어려운 샘플을 선택적으로 다시 노출합니다. 우리는 기존의 지도 학습 분류를 위해 개발된 데이터 균형 전략을 비지도 복원 환경에서 평가합니다. 이러한 접근 방식의 한계를 고려하여, 우리의 방법은 공간적 불균형에 특화되어 모델이 통계적으로 드문 위치에 집중하도록 장려하여 기존의 방법보다 복원 일관성을 향상시킵니다. 우리는 제어된 공간적 불균형 조건을 가진 시뮬레이션 데이터 세트와 실제 물리, 생물 및 천문 분야의 다양한 데이터 세트에서 검증했습니다. 우리의 방법은 다양한 복원 지표에서 기존 방법보다 우수한 성능을 보이며, 특히 공간적 불균형 분포에서 뛰어난 성능을 보입니다. 이러한 결과는 배치 내 데이터 표현의 중요성을 강조하며, 비지도 이미지 복원에서 희귀 샘플의 중요성을 보여줍니다. 우리는 모든 코드 및 관련 데이터를 공개할 예정입니다.

Original Abstract

Autoencoders can be challenged by spatially non-uniform sampling of image content. This is common in medical imaging, biology, and physics, where informative patterns occur rarely at specific image coordinates, as background dominates these locations in most samples, biasing reconstructions toward the majority appearance. In practice, autoencoders are biased toward dominant patterns resulting in the loss of fine-grained detail and causing blurred reconstructions for rare spatial inputs especially under spatial data imbalance. We address spatial imbalance by two complementary components: (i) self-entropy-based loss that upweights statistically uncommon spatial locations and (ii) Sample Propagation, a replay mechanism that selectively re-exposes the model to hard to reconstruct samples across batches during training. We benchmark existing data balancing strategies, originally developed for supervised classification, in the unsupervised reconstruction setting. Drawing on the limitations of these approaches, our method specifically targets spatial imbalance by encouraging models to focus on statistically rare locations, improving reconstruction consistency compared to existing baselines. We validate in a simulated dataset with controlled spatial imbalance conditions, and in three, uncontrolled, diverse real-world datasets spanning physical, biological, and astronomical domains. Our approach outperforms baselines on various reconstruction metrics, particularly under spatial imbalance distributions. These results highlight the importance of data representation in a batch and emphasize rare samples in unsupervised image reconstruction. We will make all code and related data available.

0 Citations
0 Influential
3.5 Altmetric
17.5 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!