2601.22904v1 Jan 30, 2026 cs.CV

DINO-SAE: DINO 기반 구면 오토인코더를 이용한 고품질 이미지 복원 및 생성

DINO-SAE: DINO Spherical Autoencoder for High-Fidelity Image Reconstruction and Generation

J. Ye
J. Ye
Citations: 87
h-index: 4
Hun Chang
Hun Chang
Citations: 17
h-index: 2
Byunghee Cha
Byunghee Cha
Citations: 18
h-index: 2

최근 연구에서는 사전 학습된 비전 기반 모델(VFMs)인 DINO를 활용하여 생성적 오토인코더를 구축하고, 뛰어난 생성 성능을 보이는 것을 확인했습니다. 그러나 기존 방법들은 고주파 상세 정보 손실로 인해 복원 품질이 제한되는 경우가 많습니다. 본 연구에서는 의미 표현과 픽셀 수준의 복원을 연결하는 프레임워크인 DINO 구면 오토인코더(DINO-SAE)를 제안합니다. 우리의 핵심 아이디어는 대조 학습 표현에서 의미 정보가 주로 특징 벡터의 방향에 인코딩된다는 점이며, 엄격한 크기 일치를 강제하면 인코더가 미세한 디테일을 보존하는 것을 방해할 수 있다는 것입니다. 이를 해결하기 위해, 우리는 로컬 구조 및 텍스처 보존을 향상시키는 계층적 컨볼루션 패치 임베딩 모듈과, 의미 일관성을 유지하면서 디테일 보존을 위한 유연한 특징 크기를 허용하는 코사인 유사성 정렬 목적 함수를 도입했습니다. 또한, 자기 지도 학습 기반의 기반 모델 표현이 본질적으로 초구(hypersphere) 위에 존재한다는 것을 활용하여, 리만 흐름 매칭(Riemannian Flow Matching)을 사용하여 확산 트랜스포머(DiT)를 이 구면 잠재 공간(spherical latent manifold)에 직접 학습시켰습니다. ImageNet-1K 데이터셋에 대한 실험 결과, 우리의 접근 방식은 최고 수준의 복원 품질을 달성했으며, 0.37의 rFID 및 26.2 dB의 PSNR을 기록하면서, 사전 학습된 VFM과의 강력한 의미 일관성을 유지했습니다. 특히, 리만 흐름 매칭 기반 DiT는 효율적인 수렴을 보여주며, 80 에포크에서 3.47의 gFID를 달성했습니다.

Original Abstract

Recent studies have explored using pretrained Vision Foundation Models (VFMs) such as DINO for generative autoencoders, showing strong generative performance. Unfortunately, existing approaches often suffer from limited reconstruction fidelity due to the loss of high-frequency details. In this work, we present the DINO Spherical Autoencoder (DINO-SAE), a framework that bridges semantic representation and pixel-level reconstruction. Our key insight is that semantic information in contrastive representations is primarily encoded in the direction of feature vectors, while forcing strict magnitude matching can hinder the encoder from preserving fine-grained details. To address this, we introduce Hierarchical Convolutional Patch Embedding module that enhances local structure and texture preservation, and Cosine Similarity Alignment objective that enforces semantic consistency while allowing flexible feature magnitudes for detail retention. Furthermore, leveraging the observation that SSL-based foundation model representations intrinsically lie on a hypersphere, we employ Riemannian Flow Matching to train a Diffusion Transformer (DiT) directly on this spherical latent manifold. Experiments on ImageNet-1K demonstrate that our approach achieves state-of-the-art reconstruction quality, reaching 0.37 rFID and 26.2 dB PSNR, while maintaining strong semantic alignment to the pretrained VFM. Notably, our Riemannian Flow Matching-based DiT exhibits efficient convergence, achieving a gFID of 3.47 at 80 epochs.

3 Citations
0 Influential
2 Altmetric
13.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!