Transformer-Driven Drug Discovery: Comparative Evaluation of ChemBERT and BioBERT in Molecular Activity Prediction
DOI:
https://doi.org/10.70917/ijcisim-2026-3485Keywords:
ChemBERT, BioBERT, Transformer Models, Molecular Activity Prediction, SMILES, Drug Discovery, Machine LearningAbstract
This paper presents a comparative analysis of two transformers-based language models, that is, ChemBERT and BioBERT, in the binary molecular activity prediction task. ChemBERT is a pretrained model that is built over large-scale chemical SMILES corpora that is supposed to replicate structural and syntax properties of molecular representations. BioBERT is, however, trained on biomedical text and trained on natural language processing tasks in life sciences. We test the two models on the provided dataset of molecular activities and find out what one would be more appropriate in predictive modeling with the help of SMILES. The experiment outcomes have revealed that ChemBERT is significantly more efficient, as compared with BioBERT with ROC-AUC of 0.8969 and accuracy of 0.9145 and 0.8391 of BioBERT. These findings prove the effectiveness of domain pretraining of chemical representation learning and the efficiency of chemics-aware transformer-based learning to predict molecular properties and predict computational drug discovery.