ProteinJEPA: 잠재 변수 예측이 단백질 언어 모델을 보완하는 방법
ProteinJEPA: Latent prediction complements protein language models
단백질 언어 모델은 주로 마스크 언어 모델링(MLM)을 사용하여 마스크된 위치의 아미노산 식별자를 예측하는 방식으로 훈련됩니다. 본 연구에서는 잠재 공간 예측이 동일한 시간 제약 조건 하에서 이러한 토큰 수준의 목표를 어떻게 보완할 수 있는지 조사했습니다. 35~150만 개의 파라미터를 가진 사전 훈련된 및 임의 초기화된 단백질 서열 인코더를 사용하여, 최적의 단백질-JEPA 설계는 모든 위치에 대한 잠재 변수 예측이 아니라 다음과 같은 변형된 방식임을 확인했습니다. 즉, 마스크된 위치에서만 잠재 변수를 예측하고 MLM의 교차 엔트로피 손실을 유지하는 것입니다. 이 방법을 '마스크된 위치 MLM + JEPA'라고 명명했습니다. 16개의 다운스트림 작업(15개의 동결된 선형 프로브 및 SCOPe-40의 제로샷 폴드 검색) 세트에서, 동일한 시간 제약 조건 하에서 이 방법은 MLM만 사용하는 방법에 비해 더 많은 작업을 성공적으로 수행했습니다. 구체적으로, 사전 훈련된 ESM2-35M 모델에서는 10승 3패 3무(W/L/T), ESM2-150M 모델에서는 11승 2패 3무, 그리고 처음부터 훈련할 경우에는 결과가 혼합되었습니다(6승 8패 2무). 여러 모델에서 16개의 작업 중 11개 작업에서 성능 향상이 나타났으며, 여기에는 안정성, β-락타마제 적합성, 변이 효과, 고유한 무작위성, 원격 유사성, 효소 분류 및 SCOPe-40 폴드 검색이 포함됩니다. 형광(TAPE) 및 펩타이드-HLA 결합 작업에서는 패배가 더 많았습니다. 모든 위치에 대한 MLM + JEPA는 전체적으로는 MLM만 사용하는 것과 비슷하지만, 마스크된 위치에서 얻을 수 있는 성능 향상을 재현하지 못합니다. JEPA만 사용하는 경우(MLM 미사용) 거의 모든 실험에서 성능이 저하되었습니다. 결론적으로, JEPA는 MLM과 결합될 때 경쟁력이 있으며, 동일한 시간 제약 조건 하에서도 사전 훈련 및 추가 훈련에서 순수한 MLM보다 더 나은 성능을 보일 수 있습니다.
Protein language models are trained primarily with masked language modeling (MLM), which predicts amino-acid identities at masked positions. We ask whether latent-space prediction can complement these token-level objectives under matched wall-clock budget. Across pretrained and random-init protein sequence encoders at 35--150M parameters, we find that the best protein-JEPA design is not all-position latent prediction but a variant: predicting latent targets only at masked positions, and retaining the MLM cross-entropy. We call this recipe masked-position MLM+JEPA. On a 16-task downstream suite (15 frozen linear probes plus SCOPe-40 zero-shot fold retrieval), under matched wall-clock budgets, this recipe wins more tasks than it loses against MLM-only continuation: 10 wins / 3 losses / 3 ties (hereafter W/L/T) on pretrained ESM2-35M, 11/2/3 on ESM2-150M while results in pretraining from scratch are mixed (6/8/2). Gains are seen for multiple models on 11 of 16 tasks, including stability, \b{eta}β\b{eta}-lactamase fitness, variant effect, intrinsic disorder, remote homology, enzyme classification, and SCOPe-40 fold retrieval. Tasks with more losses than wins are Fluorescence (TAPE) and Peptide-HLA Binding. All-position MLM+JEPA matches MLM-only overall but does not reproduce the masked-position gains. JEPA-only (no MLM) collapses in nearly every experiment. We conclude that JEPA, when combined with MLM, is competitive and can outperform pure MLM in pretraining and continued training, even under matched wall-clock budgets.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.