2608.05928v1 Aug 06, 2026 cs.LG

BioM-JEPA: 단일 세포 내 그래프로 연결된 유전자 블록의 공동 임베딩 예측

BioM-JEPA: joint-embedding prediction of graph-connected gene blocks in single cells

Ziyun Zhang
Ziyun Zhang
Citations: 0
h-index: 0
Zelin Zang
Zelin Zang
Citations: 550
h-index: 14
Zhen Lei
Zhen Lei
Citations: 18
h-index: 3
Stan Z. Li
Stan Z. Li
Citations: 75
h-index: 5
Yuhao Wang
Yuhao Wang
Citations: 2
h-index: 1

단일 세포 트랜스크립톰 데이터는 조화로운 생물학적 프로그램의 희소한 관찰 결과이지만, 대부분의 자기 지도 학습 모델은 개별 유전자를 재구성하는 방식으로 학습합니다. 본 연구에서는 단백질 상호작용 및 코퍼스 기반 공존 발현 증거로 정의된 그래프로 연결된 유전자 블록의 집계 표현을 예측하는 공동 임베딩 예측 아키텍처인 BioM-JEPA를 제시합니다. 학생 네트워크는 세포 내 나머지 유전자를 사용하여 각 대상 블록 표현을 추론하고, 천천히 업데이트되는 교사 네트워크는 전체 관찰 유전자 세트에서 해당 대상을 제공합니다. 보고된 추출 절차 하에서, 블록 수준 예측은 테스트된 진단에서 토큰 예측, 임의 블록 및 재구성 제어보다 더 높은 효과적인 순위와 검출된 유전자 수에 대한 약한 연관성을 가진 임베딩을 생성했습니다. CellBench 작업 전반에 걸쳐, 동결된 BioM-JEPA 임베딩은 발현, 경로 및 이웃 정보를 유지했으며 평가된 모델 중에서 가장 낮은 집계 교란-반응 오류를 달성했습니다. 표현 진단 결과는 표준적인 췌장 프로그램 및 유전자 변형 간의 구성적 관계와 일관성을 보였습니다. 선형 어텐션은 이차적인 유전자-유전자 어텐션 행렬을 구축하지 않으며, hPancreas 실험에서 배치 크기 8로 한 에포크 동안 BioM-JEPA는 scFoundation보다 5.75배 더 높은 미세 조정 처리량과 3.76배 더 높은 보류 임베딩 처리량을 제공했습니다. 이러한 결과들을 종합적으로 고려할 때, 그래프로 연결된 유전자 블록은 단일 세포 생물학에서 JEPA 스타일의 표현 학습을 위한 유용한 예측 단위임을 알 수 있습니다.

Original Abstract

Single-cell transcriptomes are sparse observations of coordinated biological programmes, yet most self-supervised models learn by reconstructing individual genes. Here we present BioM-JEPA, a joint-embedding predictive architecture that instead predicts aggregate representations of graph-connected gene blocks defined by protein-association and corpus-derived coexpression evidence. A student network infers each target-block representation from the remaining genes in a cell, while a slowly updated teacher supplies the corresponding target from the full observed gene set. Under the reported extraction procedure, block-level prediction produced embeddings with higher effective rank and weaker association with detected-gene depth in the tested diagnostics than token-prediction, random-block and reconstruction controls. Across CellBench tasks, frozen BioM-JEPA embeddings retained expression, pathway and neighbourhood information and achieved the lowest aggregate perturbation-response error among the evaluated models. Representation diagnostics were also consistent with canonical pancreatic programmes and compositional relationships between genetic perturbations. Linear attention avoids constructing a quadratic gene-by-gene attention matrix; in a matched one-epoch hPancreas experiment at batch size 8, BioM-JEPA provided 5.75-fold higher fine-tuning throughput and 3.76-fold higher held-out embedding throughput than scFoundation. Together, these results support graph-connected gene blocks as useful prediction units for JEPA-style representation learning in single-cell biology.

0 Citations
0 Influential
7 Altmetric
35.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!