PepSpecBench: 펩타이드 탠덤 질량 분석 예측을 위한 통합 평가 벤치마크
PepSpecBench: A Unified Evaluation Benchmark for Peptide Tandem Mass Spectrometry Prediction
탠덤 질량 분석법은 복잡한 생물학적 샘플에서 단백질을 식별하고 정량화하는 데 유용한 고처리량 프레임워크를 제공합니다. 계산 단백질체학에서 펩타이드 MS/MS 스펙트럼을 예측하는 것은 대규모 펩타이드 식별 및 정량화와 같은 후속 응용을 가능하게 하는 중요한 과제입니다. 딥러닝 아키텍처는 예측 정확도를 크게 향상시켰지만, 세 가지 평가상의 어려움이 이 분야의 진정한 발전을 가리고 있습니다. 첫째, 일관성 없는 데이터 전처리 및 호환되지 않는 모델 출력 공간은 공정한 모델 비교를 방해합니다. 둘째, 잘못된 데이터 분할 전략은 숨겨진 시퀀스 누수를 초래하고 보고된 성능을 과대평가할 수 있습니다. 셋째, 기존 평가에서는 포괄적인 종 간 벤치마킹과 모델의 견고성을 영향력 있는 실험 조건에 대해 체계적으로 평가하는 것이 부족합니다. 이러한 문제점을 해결하기 위해 펩타이드 MS/MS 스펙트럼 예측을 위한 통합 벤치마크인 PepSpecBench를 제안합니다. PepSpecBench는 상호 보완적인 공개 데이터 세트를 사용하여 데이터 전처리를 표준화하고, 시퀀스 누수를 제거하기 위한 엄격한 백본 분리 분할 전략을 적용하며, 공유된 조각 이온 표현 공간 내에서 다양한 아키텍처를 평가합니다. 또한, 포괄적인 다종 평가 세트와 물리적으로 근거한 메타데이터 교란 테스트를 도입하여 모델의 견고성과 장비 인지 능력을 평가합니다. 우리는 여섯 가지 대표 모델에서 이전에 인식되지 못했던 성능 차이와 견고성 한계를 발견했으며, 이는 향후 모델 설계, 평가 및 실제 배포를 위한 실질적인 통찰력을 제공합니다.
Tandem mass spectrometry provides a high-throughput framework for identifying and quantifying proteins in complex biological samples. In computational proteomics, predicting peptide MS/MS spectra is a critical task, enabling downstream applications such as large-scale peptide identification and quantification. While deep learning architectures have substantially improved prediction accuracy, three evaluation challenges obscure the true progress of the field. First, inconsistent data preprocessing and incompatible model output spaces hinder fair model comparison. Second, flawed data splitting strategies can permit hidden sequence leakage and inflate reported performance. Third, existing evaluations typically lack comprehensive cross-species benchmarking and systematic assessment of model robustness to influential experimental conditions. To address these challenges, we propose PepSpecBench, a unified benchmark for peptide MS/MS spectrum prediction. PepSpecBench standardizes data preprocessing across complementary public datasets, enforces a strict backbone-disjoint splitting strategy to eliminate sequence leakage, and evaluates diverse architectures within a shared fragment-ion representation space. It further introduces a comprehensive multi-species evaluation suite and physically grounded metadata perturbation probes to assess model robustness and instrument awareness. We uncover previously unrecognized performance discrepancies and robustness limitations across six representative models, providing actionable insights for future model design, evaluation and practical deployment.
No Analysis Report Yet
This paper hasn't been analyzed by Gemini yet.
Log in to request an AI analysis.