Attention-Guided Vision-Language Model for Automated Radiology Report Generation

Authors

  • Dr. P. Dayaker School of Computer Science and Engineering, Malla Reddy (MR) Deemed to be University, Hyderabad, India.
  • M. Vignesh Department of Artificial Intelligence and Data Science, Karpagam Institute of Technology, Coimbatore, Tamil Nadu, India.
  • Dr. Inamul Hussain R. Z. Department of Computer Science and Engineering, C. Abdul Hakeem College of Engineering and Technology, Melvisharam, Tamil Nadu, India.
  • R. Rajkumar Department of Electronics and Communication Engineering, Vel Tech Rangarajan Dr. Sagunthala R&D Institute of Science and Technology, Avadi, Chennai, Tamil Nadu 600062, India.
  • Aruna P. Department of Computer Science and Engineering, Nehru Institute of Technology, Jawahar Gardens, Kaliapuram, Coimbatore, Tamil Nadu, India.
  • Jeevanantham G. Department of Computer Science and Engineering, Nehru Institute of Engineering and Technology (Autonomous), Coimbatore, Tamil Nadu, India.

DOI:

https://doi.org/10.70917/ijcisim-2026-5290

Keywords:

Automated Radiology Report Generation, Vision-Language Model, Attention Mechanism, Medical Image Analysis, Transformer, Chest Radiography, Multimodal Learning, Clinical Natural Language Generation

Abstract

Automated radiology report generation (ARRG) has emerged as a promising application of artificial intelligence for reducing radiologists’ documentation workload and improving the consistency of clinical reporting. However, conventional image-to-text models often struggle to capture subtle abnormalities, establish meaningful associations between localized visual findings and clinical terminology, and generate diagnostically relevant descriptions. This study proposes an Attention-Guided Vision-Language Model (AG-VLM) for automated radiology report generation that integrates multi-scale visual feature extraction, spatial attention-guided abnormality localization, cross-modal vision-language alignment, and an attention-aware Transformer-based report decoder. The proposed framework selectively emphasizes clinically significant image regions while suppressing redundant background information, thereby strengthening the correspondence between radiographic findings and generated textual descriptions. Experiments were conducted using chest radiograph–report pairs, with performance evaluated using standard natural-language-generation and clinical-consistency measures. The proposed AG-VLM achieved a BLEU-1 score of 0.521, BLEU-2 of 0.387, BLEU-3 of 0.301, BLEU-4 of 0.243, METEOR of 0.286, ROUGE-L of 0.418, and CIDEr of 0.472. For clinical content preservation, the framework obtained a clinical precision of 0.861, recall of 0.842, and F1-score of 0.851. The attention-guided architecture also achieved an abnormality localization accuracy of 91.7% and an overall clinical finding accuracy of 92.4%. Compared with the selected baseline vision-language report-generation model, AG-VLM improved BLEU-4 by 12.5%, METEOR by 9.6%, ROUGE-L by 8.3%, and clinical F1-score by 7.9%. These results indicate that explicit attention-guided visual reasoning combined with cross-modal semantic alignment can generate more accurate, clinically coherent, and contextually relevant radiology reports. The proposed framework therefore provides a scalable foundation for computer-assisted radiology reporting while retaining the need for radiologist verification before clinical use.

Downloads

Download data is not yet available.

Downloads

Published

2026-08-30

How to Cite

Dr. P. Dayaker, M. Vignesh, Dr. Inamul Hussain R. Z., R. Rajkumar, Aruna P., & Jeevanantham G. (2026). Attention-Guided Vision-Language Model for Automated Radiology Report Generation. International Journal of Computer Information Systems and Industrial Management Applications, 18(21s), 200–216. https://doi.org/10.70917/ijcisim-2026-5290

Issue

Section

Original Articles