2602.22822v1 Feb 26, 2026 cs.AI

FlexMS: 대사체학 분야의 심층 학습 기반 질량 스펙트럼 예측 도구 벤치마킹을 위한 유연한 프레임워크

FlexMS is a flexible framework for benchmarking deep learning-based mass spectrum prediction tools in metabolomics

Yifan Li
Yifan Li
Citations: 4
h-index: 1
Jie Yang
Jie Yang
Citations: 26
h-index: 4
Panlong Liu
Panlong Liu
Citations: 3
h-index: 1
Yun-Yan Zhong
Yun-Yan Zhong
Citations: 11
h-index: 2
Yixuan Tang
Yixuan Tang
Citations: 35
h-index: 3
Jun Xia
Jun Xia
Citations: 23
h-index: 2
Zhiwen Yang
Zhiwen Yang
Citations: 24
h-index: 1
Zanyi Wang
Zanyi Wang
Citations: 3
h-index: 1

화학 분자의 식별 및 특성 예측은 신약 개발 및 재료 과학 발전의 핵심이며, 탠덤 질량 분석 기술은 질량-전하 비율 피크 형태로 유용한 분해 정보를 제공합니다. 그러나 실험적 스펙트럼의 부족은 각 분자 식별에 어려움을 야기하며, 이를 해결하기 위한 예측 방법론 개발의 필요성을 강조합니다. 심층 학습 모델은 분자 구조 스펙트럼 예측에 유망하지만, 방법론의 다양성과 잘 정의된 벤치마크의 부족으로 인해 전체적인 평가가 어렵습니다. 이에, 우리는 질량 스펙트럼 예측 모델의 다양한 아키텍처를 구축하고 평가하기 위한 벤치마크 프레임워크 FlexMS를 개발했습니다. FlexMS는 사용하기 쉬운 유연성을 제공하여, 다양한 모델 아키텍처 조합을 동적으로 구성하고, 다양한 지표를 사용하여 전처리된 공개 데이터셋에 대한 성능을 평가할 수 있습니다. 본 논문에서는 데이터셋의 구조적 다양성, 학습률 및 데이터 희소성과 같은 하이퍼파라미터, 사전 학습 효과, 메타데이터 제거 설정, 그리고 교차 도메인 전이 학습 분석을 포함하여 성능에 영향을 미치는 요인에 대한 통찰력을 제공합니다. 이를 통해 적합한 모델 선택에 대한 실질적인 지침을 제공합니다. 또한, 검색 벤치마크는 실제 식별 시나리오를 시뮬레이션하고 예측된 스펙트럼을 기반으로 잠재적인 일치 항목을 평가합니다.

Original Abstract

The identification and property prediction of chemical molecules is of central importance in the advancement of drug discovery and material science, where the tandem mass spectrometry technology gives valuable fragmentation cues in the form of mass-to-charge ratio peaks. However, the lack of experimental spectra hinders the attachment of each molecular identification, and thus urges the establishment of prediction approaches for computational models. Deep learning models appear promising for predicting molecular structure spectra, but overall assessment remains challenging as a result of the heterogeneity in methods and the lack of well-defined benchmarks. To address this, our contribution is the creation of benchmark framework FlexMS for constructing and evaluating diverse model architectures in mass spectrum prediction. With its easy-to-use flexibility, FlexMS supports the dynamic construction of numerous distinct combinations of model architectures, while assessing their performance on preprocessed public datasets using different metrics. In this paper, we provide insights into factors influencing performance, including the structural diversity of datasets, hyperparameters like learning rate and data sparsity, pretraining effects, metadata ablation settings and cross-domain transfer learning analysis. This provides practical guidance in choosing suitable models. Moreover, retrieval benchmarks simulate practical identification scenarios and score potential matches based on predicted spectra.

0 Citations
0 Influential
2 Altmetric
10.0 Score
Original PDF

No Analysis Report Yet

This paper hasn't been analyzed by Gemini yet.

Log in to request an AI analysis.

댓글

댓글을 작성하려면 로그인하세요.

아직 댓글이 없습니다. 첫 번째 댓글을 남겨보세요!