TMT-SpecGen: A Computational Utility for Boosting Peptide Identification via TMT Spectral Library Augmentation
DOI:
https://doi.org/10.70917/ijcisim-2026-3745Keywords:
Deep learning, TMT-label, spectral library, annotated spectra, spectrum prediction, peptide identificationAbstract
Multiplexed Tandem Mass Tag (TMT) proteomics is a well-established technique widely employed for large-scale peptide quantification and identification. However, traditional sequence database searches for peptide identification are obstructed by the effects of TMT labeling. Spectral library searches offer improved detection capability (true positives) and selectivity (true negatives) for peptide identification in TMT-labeled mass spectrometry data. In this work, we aim to increase peptide identification by enhancing spectral library. The fundamental principle underlying the proposed work is that a peptide's tandem MS fragmentation pattern, under consistent conditions, is reproducible. This reproducibility enables the identification of unknown spectra acquired under similar conditions through spectral matching. To achieve this, we introduce TMT-SpecGen (Tandem Mass Tag- Spectrum Generator), a biologically inspired utility that generates annotated spectra using a deep learning model. It draws concepts from the way the human brain processes information, learns from experience, and adapts to new data. The model is trained on labeled spectra obtained from clustering and consensus generation. Utilizing consensus spectra for training reduces training time by 40% while improving accuracy of spectrum prediction by 5%. From 163,615 predictions, we constructed a high-quality simulated TMT spectral library containing 96,631 robust peptides derived from millions of peptide-spectrum matches. The simulated library demonstrates its efficacy in enhancing peptide identification, as 14% of previously unidentified spectra, when searched in the library, were successfully identified.