2604.16570v1 Apr 17, 2026 cs.LG

잃어버린 DNA 서열 사전 훈련 탐색

In Search of Lost DNA Sequence Pretraining

Zhijiang Tang
Zhijiang Tang
Citations: 12
h-index: 2
Jianqiang Huang
Jianqiang Huang
Citations: 7
h-index: 2
Yuhua Zheng
Yuhua Zheng
Citations: 9
h-index: 1
Jiaxin Qi
Jiaxin Qi
Citations: 6
h-index: 2
Yan Cui
Yan Cui
Citations: 4
h-index: 1
Jinli Ou
Jinli Ou
Citations: 1
h-index: 1

DNA 서열 인코딩은 유전자 기능 예측, 단백질 합성 및 다양한 생물학적 응용 분야에서 근본적인 역할을 수행합니다. 대규모 DNA 서열 사전 훈련을 통해 상당한 발전이 이루어졌음에도 불구하고, 기존 연구는 주로 사전 훈련의 규모와 맞춤형 평가 데이터셋에 초점을 맞추고 사전 훈련 패러다임의 필수적인 요소들을 간과했습니다. 본 논문에서는 DNA 사전 훈련에서 간과되어 왔던 세 가지 중요한 문제를 밝혀냅니다. 즉, 부적절한 다운스트림 데이터셋, 이웃 마스킹 전략의 근본적인 결함, 그리고 어휘에 대한 상세한 논의 부족입니다. 따라서, 우리는 포괄적인 조사를 수행하고 평가 데이터셋 선택 기준, 작업 설계 지침, 그리고 심층적인 어휘 분석을 포함한 원칙적인 지침을 제안합니다. 광범위한 실험을 통해 우리가 지적한 문제의 중요성을 검증하고, 우리의 제안에 대한 근거를 뒷받침합니다. 마지막으로, 우리는 DNA 사전 훈련 방법의 재현 가능하고 엄격한 벤치마킹을 가능하게 하는 표준화된 테스트 환경을 소개하여, 유전체 기반 모델 개발을 촉진하고자 합니다.

Original Abstract

DNA sequence encoding is fundamental to gene function prediction, protein synthesis, and diverse downstream biological tasks. Despite the substantial progress achieved by large-scale DNA sequence pretraining, existing studies have overwhelmingly emphasized pretraining scale and custom downstream evaluation datasets, while neglecting some essential components of the pretraining paradigm. In this paper, we reveal three critical yet heretofore overlooked problems in DNA pretraining: inappropriate downstream datasets, inherent flaws in the neighbor-masking strategy, and the lack of detailed discussion on vocabulary. Therefore, we undertake comprehensive investigations and propose principled guidelines, including selection criteria for evaluation datasets, guiding task design, and in-depth vocabulary analysis. Extensive experiments validate the significance of our identified problems and support the rationale behind our recommendations. Finally, we introduce a standardized testbed that enables reproducible and rigorous benchmarking of DNA pretraining methods to advance the development of genomic foundation models.

0 Citations
0 Influential
1 Altmetric
5.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!