대규모 사전 학습을 위한 공간 트랜스크립토믹스: 이미지 기반 접근 방식
Spatial Transcriptomics as Images for Large-Scale Pretraining
공간 트랜스크립토믹스(ST)는 조직 슬라이드 상의 특정 위치에서 수천 개의 유전자 발현 값을 정밀한 좌표와 함께 측정하여, 임상 및 병리학 연구에 필수적인 공간적 맥락을 보존합니다. 시퀀싱 기술의 발전과 데이터 처리 능력 향상으로 인해, ST 데이터의 양이 증가하면서 대규모 ST 사전 학습의 필요성이 커지고 있습니다. 그러나 사전 학습의 기본 단위, 즉 하나의 학습 샘플이 무엇을 구성하는지에 대한 명확한 정의는 아직 부족합니다. 기존의 방법은 크게 두 가지로 나뉩니다. (1) 각 위치를 독립적인 샘플로 취급하는 방법은 공간적 의존성을 무시하고 ST를 단일 세포 트랜스크립토믹스로 단순화합니다. (2) 전체 슬라이드를 하나의 샘플로 취급하는 방법은 입력 데이터의 크기가 지나치게 커지고 학습 샘플의 수가 현저히 감소하여 효과적인 사전 학습을 어렵게 만듭니다. 이러한 문제점을 해결하기 위해, 본 연구에서는 공간 트랜스크립토믹스 데이터를 편집 가능한 이미지로 취급하는 새로운 방법을 제안합니다. 구체적으로, 원본 슬라이드에서 특정 크기의 영역을 추출하여 고정된 공간 크기의 다중 채널 이미지 표현을 정의함으로써, 공간적 맥락을 유지하면서 학습 샘플의 수를 크게 늘립니다. 또한, 채널 차원을 따라 유전자 부분 집합 선택 규칙을 정의하여 입력 데이터의 차원을 조절하고 사전 학습의 안정성을 향상시킵니다. 광범위한 실험 결과, 제안하는 이미지 기반 ST 데이터 구축 방법은 기존의 사전 학습 방법보다 일관되게 우수한 성능을 보였습니다. 추가적인 분석을 통해, 공간 패칭과 채널 설계 모두가 필수적임을 확인했으며, 이를 통해 ST 데이터를 체계적으로 구성하고 대규모 사전 학습을 가능하게 하는 실용적인 방법을 제시합니다.
Spatial Transcriptomics (ST) profiles thousands of gene expression values at discrete spots with precise coordinates on tissue sections, preserving spatial context essential for clinical and pathological studies. With rising sequencing throughput and advancing platforms, the expanding data volumes motivate large-scale ST pretraining. However, the fundamental unit for pretraining, i.e., what constitutes a single training sample, remains ill-posed. Existing choices fall into two camps: (1) treating each spot as an independent sample, which discards spatial dependencies and collapses ST into single-cell transcriptomics; and (2) treating an entire slide as a single sample, which produces prohibitively large inputs and drastically fewer training examples, undermining effective pretraining. To address this gap, we propose treating spatial transcriptomics as croppable images. Specifically, we define a multi-channel image representation with fixed spatial size by cropping patches from raw slides, thereby preserving spatial context while substantially increasing the number of training samples. Along the channel dimension, we define gene subset selection rules to control input dimensionality and improve pretraining stability. Extensive experiments show that the proposed image-like dataset construction for ST pretraining consistently improves downstream performance, outperforming conventional pretraining schemes. Ablation studies verify that both spatial patching and channel design are necessary, establishing a unified, practical paradigm for organizing ST data and enabling large-scale pretraining.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.