2608.08366v1 Aug 08, 2026 cs.CV

VOICE: 영상-유전체 기반 모델 - 현장 단일 세포 유전자 발현의 직접 예측 및 검색 기반 예측 통합

VOICE: A Vision-Omics Foundation Model Integrating Direct and Retrieval-Based Prediction of In-situ Single-Cell Gene Expression

Xin Luo
Xin Luo
Citations: 32
h-index: 4
Yicheng Tao
Yicheng Tao
Citations: 75
h-index: 4
Haoxuan Zeng
Haoxuan Zeng
Citations: 38
h-index: 2
Suyuan Wang
Suyuan Wang
Citations: 44
h-index: 2
Chenzi Ouyang
Chenzi Ouyang
Citations: 0
h-index: 0
Meiqing Zhu
Meiqing Zhu
Citations: 0
h-index: 0
Kai Liu
Kai Liu
Citations: 312
h-index: 3
Shuibing Chen
Shuibing Chen
Citations: 32
h-index: 3
Jie Liu
Jie Liu
Citations: 6
h-index: 1

공간 트랜스크립토믹스는 단일 세포 수준에서 유전자 발현을 분석할 수 있지만, 비용이 많이 들고, 수백에서 수천 개의 유전자에 한정되며, 소수의 샘플에만 적용 가능합니다. 반면, H&E 이미징은 저렴하며 대규모로 정기적으로 획득됩니다. 따라서, 형태학적 특징으로부터 직접 단일 세포 발현을 예측하는 것은 분자 분석을 대량의 조직 자료에 적용할 수 있는 실용적인 방법입니다. 이에 우리는 H&E 이미지와 연계된 Xenium 데이터를 사용하여 단일 세포 유전자 발현을 예측하는 다중 모드 기반 모델인 VOICE를 제시합니다. VOICE는 먼저 병리학 기반 모델에서 얻은 세포 중심 H&E 형태학 정보를, 2300만 개의 세포에 대한 대비 학습으로 훈련된 트랜스크립톰 기반 모델의 단일 세포 발현 임베딩과 정렬합니다. 그런 다음, 두 가지 방식으로 유전자 발현을 예측합니다. 한 가지 방식은 형태학적 특징으로부터 직접적으로 유전자 발현을 회귀시키는 것이고, 다른 방식은 유사한 기준 세포로부터 측정된 유전자 발현 값을 검색하여 형태학적 신호가 없는 유전자를 복원하는 것입니다. 유전자별로 형태학적 예측 가능성이 다르기 때문에, VOICE는 각 유전자마다 가중치를 부여하여 두 가지 방식을 통합합니다. 훈련 후, VOICE는 Xenium의 보류된 환자, 슬라이드 및 부분적으로 중복되는 유전자 패널에 대해 일반화 성능을 보이며, 일관되게 기존의 단일 세포 발현 예측 방법보다 7가지 지표에서 더 우수한 성능을 나타냅니다.

Original Abstract

Spatial transcriptomics can resolve gene expression at single-cell resolution, but it is costly, limited to targeted panels of a few hundred to a few thousand genes, and applicable to only a small number of samples. H&E imaging, by contrast, is cheap and collected routinely at scale. This makes predicting single-cell expression directly from morphology a practical way to bring molecular analysis to large tissue archives. We therefore present VOICE, a multimodal foundation model that predicts single-cell gene expression from H&E images using paired Xenium data. VOICE first aligns cell centered H&E morphology from a pathology foundation model with single-cell expression embeddings from a transcriptome foundation model, trained using contrastive learning over 23 million cells. Next it predicts expression through two branches. One branch directly regresses expression from morphology. The other branch retrieves measured expression from similar reference cells, recovering genes that do not have morphological signal. Because genes vary in morphological predictability, VOICE fuses the two branches with a per-gene weight. After training, VOICE generalizes to heldout patients, slides, and partially overlapping gene panels from Xenium, and it consistently outperforms prior single-cell expression prediction methods on seven metrics.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!