BioM-JEPA: 단일 세포 내 그래프로 연결된 유전자 블록의 공동 임베딩 예측
BioM-JEPA: joint-embedding prediction of graph-connected gene blocks in single cells
단일 세포 트랜스크립톰 데이터는 조화로운 생물학적 프로그램의 희소한 관찰 결과이지만, 대부분의 자기 지도 학습 모델은 개별 유전자를 재구성하는 방식으로 학습합니다. 본 연구에서는 단백질 상호작용 및 코퍼스 기반 공존 발현 증거로 정의된 그래프로 연결된 유전자 블록의 집계 표현을 예측하는 공동 임베딩 예측 아키텍처인 BioM-JEPA를 제시합니다. 학생 네트워크는 세포 내 나머지 유전자를 사용하여 각 대상 블록 표현을 추론하고, 천천히 업데이트되는 교사 네트워크는 전체 관찰 유전자 세트에서 해당 대상을 제공합니다. 보고된 추출 절차 하에서, 블록 수준 예측은 테스트된 진단에서 토큰 예측, 임의 블록 및 재구성 제어보다 더 높은 효과적인 순위와 검출된 유전자 수에 대한 약한 연관성을 가진 임베딩을 생성했습니다. CellBench 작업 전반에 걸쳐, 동결된 BioM-JEPA 임베딩은 발현, 경로 및 이웃 정보를 유지했으며 평가된 모델 중에서 가장 낮은 집계 교란-반응 오류를 달성했습니다. 표현 진단 결과는 표준적인 췌장 프로그램 및 유전자 변형 간의 구성적 관계와 일관성을 보였습니다. 선형 어텐션은 이차적인 유전자-유전자 어텐션 행렬을 구축하지 않으며, hPancreas 실험에서 배치 크기 8로 한 에포크 동안 BioM-JEPA는 scFoundation보다 5.75배 더 높은 미세 조정 처리량과 3.76배 더 높은 보류 임베딩 처리량을 제공했습니다. 이러한 결과들을 종합적으로 고려할 때, 그래프로 연결된 유전자 블록은 단일 세포 생물학에서 JEPA 스타일의 표현 학습을 위한 유용한 예측 단위임을 알 수 있습니다.
Single-cell transcriptomes are sparse observations of coordinated biological programmes, yet most self-supervised models learn by reconstructing individual genes. Here we present BioM-JEPA, a joint-embedding predictive architecture that instead predicts aggregate representations of graph-connected gene blocks defined by protein-association and corpus-derived coexpression evidence. A student network infers each target-block representation from the remaining genes in a cell, while a slowly updated teacher supplies the corresponding target from the full observed gene set. Under the reported extraction procedure, block-level prediction produced embeddings with higher effective rank and weaker association with detected-gene depth in the tested diagnostics than token-prediction, random-block and reconstruction controls. Across CellBench tasks, frozen BioM-JEPA embeddings retained expression, pathway and neighbourhood information and achieved the lowest aggregate perturbation-response error among the evaluated models. Representation diagnostics were also consistent with canonical pancreatic programmes and compositional relationships between genetic perturbations. Linear attention avoids constructing a quadratic gene-by-gene attention matrix; in a matched one-epoch hPancreas experiment at batch size 8, BioM-JEPA provided 5.75-fold higher fine-tuning throughput and 3.76-fold higher held-out embedding throughput than scFoundation. Together, these results support graph-connected gene blocks as useful prediction units for JEPA-style representation learning in single-cell biology.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.