Transformer-based Discovery of Antimicrobial Peptides and Prediction of their Antibacterial Activity

Authors

  • Shuwen Pan Department of Medical Microbiology & Immunology, Faculty of Medicine, Universiti Kebangsaan Malaysia (UKM), Cheras, Kuala Lumpur, 56000, Malaysia
  • Eason Soo Department of Restorative Dentistry, Faculty of Dentistry, Universiti Kebangsaan Malaysia (UKM), Kuala Lumpur, 56000, Malaysia
  • Konken Wong Department of Medical Microbiology & Immunology, Faculty of Medicine, Universiti Kebangsaan Malaysia (UKM), Cheras, Kuala Lumpur, 56000, Malaysia

DOI:

https://doi.org/10.70917/ijcisim-2026-3793

Keywords:

antimicrobial peptides; Transformer; protein language model; self-supervised learning; deep learning; activity prediction; drug discovery

Abstract

With the worsening crisis of antimicrobial resistance, many researchers are now exploring new types of antibacterial agents, and antimicrobial peptides (AMPs) have begun to attract much attention. However, the experimental identification of AMPs in the large space of natural and synthetic sequences is still slow, expensive and labour-intensive. AMP-Transformer is a two-stage deep learning framework that combines self-supervised pre-training on large-scale unlabelled protein corpora with supervised fine-tuning on curated AMP datasets in this study. The model is a multi-layer bidirectional Transformer encoder that learns contextual residue representations via masked residue modelling, and is then fine-tuned for two coupled tasks: binary discrimination of AMPs from non-AMPs and regression of minimum inhibitory concentration (MIC) values. We used the benchmark datasets built by DBAASP v3, DRAMP 4.0 and other recently published experimental collections for evaluation. On an independent test set, AMP-Transformer had an accuracy of 95.3%, a Matthews correlation coefficient of 0.906, and an area under the receiver operating characteristic curve (AUC-ROC) of 0.986, and outperformed the support vector machine, random forest, convolutional, recurrent and hybrid baselines significantly. Ablation studies show that self-supervised pre-training and multi-head self-attention are the two main contributors, accounting for about 4% of the accuracy increase. Analysis of the learned attention maps shows that the model has independently learned the amphipathic periodicity and cationic residue enrichment characteristic of membrane-active peptides, thereby providing a degree of mechanistic interpretability that is rare among black-box predictors. A subsequent screening of metagenomic open reading frames also produced a ranked list of candidate AMPs with low sequence identity to any training examples, demonstrating the value of the framework for early-stage discovery. Based on the above results, Transformer-based protein language models are relatively stable, interpretable and scalable paradigms for AMP discovery and activity prediction, and they can help establish a practical computational pipeline that selects promising peptide candidates for experimental validation at a lower cost compared with traditional methods.

Downloads

Download data is not yet available.

Downloads

Published

2026-07-28

How to Cite

Shuwen Pan, Eason Soo, & Konken Wong. (2026). Transformer-based Discovery of Antimicrobial Peptides and Prediction of their Antibacterial Activity. International Journal of Computer Information Systems and Industrial Management Applications, 18, 14. https://doi.org/10.70917/ijcisim-2026-3793

Issue

Section

Original Articles