2602.17162v1 Feb 19, 2026 cs.AI

JEPA-DNA: 결합 임베딩 예측 아키텍처를 통한 유전체 파운데이션 모델 그라운딩

JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures

Ariel Larey
Ariel Larey
Citations: 103
h-index: 6
Elay Dahan
Elay Dahan
Citations: 19
h-index: 3
Amit Bleiweiss
Amit Bleiweiss
Citations: 69
h-index: 4
R. Kellerman
R. Kellerman
Citations: 35
h-index: 4
Guy Leib
Guy Leib
Citations: 12
h-index: 2
O. Nayshool
O. Nayshool
Citations: 175
h-index: 8
T. Zinger
T. Zinger
Citations: 160
h-index: 4
D. Dominissini
D. Dominissini
Citations: 11,933
h-index: 27
Gideon Rechavi
Gideon Rechavi
Citations: 1,270
h-index: 17
Nicole Bussola
Nicole Bussola
Citations: 28
h-index: 4
Dung Hoang
Dung Hoang
Citations: 12
h-index: 2
Nati Daniel
Nati Daniel
Citations: 17
h-index: 3
Yoli Shavit
Yoli Shavit
Citations: 886
h-index: 14
Dan Ofer
Dan Ofer
Hebrew University of Jerusalem
Citations: 2,013
h-index: 13
Simon A. Lee
Simon A. Lee
UCLA
Citations: 246
h-index: 10
Shane O’Connell
Shane O’Connell
Citations: 20
h-index: 3
Marissa Wirth
Marissa Wirth
Citations: 90
h-index: 7
A. Charney
A. Charney
Citations: 15,996
h-index: 56

유전체 파운데이션 모델(GFM)은 생명의 언어를 학습하기 위해 주로 마스크 언어 모델링(MLM) 또는 다음 토큰 예측(NTP)에 의존해 왔다. 이러한 패러다임은 국소적인 유전체 구문과 세밀한 모티프 패턴을 포착하는 데는 뛰어나지만, 더 넓은 기능적 맥락을 포착하는 데는 종종 실패하여 전역적인 생물학적 관점이 결여된 표현을 초래한다. 우리는 결합 임베딩 예측 아키텍처(JEPA)를 전통적인 생성 목적 함수와 통합한 새로운 사전 학습 프레임워크인 JEPA-DNA를 소개한다. JEPA-DNA는 CLS 토큰을 지도(supervising)하여 잠재 공간에서의 예측 목적 함수와 토큰 수준의 복원을 결합함으로써 잠재 그라운딩(latent grounding)을 도입한다. 이는 모델이 개별 뉴클레오타이드에만 집중하는 대신 마스킹된 유전체 분절의 고차원적인 기능적 임베딩을 예측하도록 강제한다. JEPA-DNA는 NTP 및 MLM 패러다임을 모두 확장하며, 처음부터(from-scratch) 학습하기 위한 독립적인 목적 함수로 사용되거나 기존 GFM을 위한 지속적 사전 학습 강화 기법으로 적용될 수 있다. 다양한 유전체 벤치마크 제품군에 걸친 평가를 통해, 우리는 JEPA-DNA가 생성 전용 베이스라인에 비해 지도 학습 및 제로샷(zero-shot) 작업에서 일관되게 우수한 성능을 달성함을 입증한다. JEPA-DNA는 보다 강력하고 생물학적 기반이 탄탄한 표현을 제공함으로써, 유전체 알파벳뿐만 아니라 서열의 기저에 있는 기능적 논리까지 이해하는 파운데이션 모델로 향하는 확장 가능한 경로를 제공한다.

Original Abstract

Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Nature". While effective at capturing local syntax, these generative paradigms prioritize token-level reconstruction over high-level functional context. We introduce JEPA-DNA, a model-agnostic continual training framework that integrates a Joint-Embedding Predictive Architecture (JEPA) with traditional generative objectives. By supervising global sequence embeddings in a latent space, JEPA-DNA forces models to predict the functional representations of masked genomic segments, shifting the learning signal from token recovery to semantic alignment. We evaluate JEPA-DNA on 17 diverse genomic benchmark tasks, demonstrating consistent gains in linear probing and zero-shot performance regardless of the underlying GFM architecture or generative objective. Our framework establishes a new state-of-the-art for GFMs, surpassing the best existing models by bridging generative precision with latent semantic grounding. Through extensive ablation studies, we further characterize the synergistic interplay between generative and latent objectives. Our code is publicly available at https://github.com/NVIDIA-Digital-Bio/JEPA-DNA.

8 Citations
0 Influential
28 Altmetric
148.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!