Attention-Guided Vision-Language Model for Automated Radiology Report Generation
DOI:
https://doi.org/10.70917/ijcisim-2026-5290Keywords:
Automated Radiology Report Generation, Vision-Language Model, Attention Mechanism, Medical Image Analysis, Transformer, Chest Radiography, Multimodal Learning, Clinical Natural Language GenerationAbstract
Automated radiology report generation (ARRG) has emerged as a promising application of artificial intelligence for reducing radiologists’ documentation workload and improving the consistency of clinical reporting. However, conventional image-to-text models often struggle to capture subtle abnormalities, establish meaningful associations between localized visual findings and clinical terminology, and generate diagnostically relevant descriptions. This study proposes an Attention-Guided Vision-Language Model (AG-VLM) for automated radiology report generation that integrates multi-scale visual feature extraction, spatial attention-guided abnormality localization, cross-modal vision-language alignment, and an attention-aware Transformer-based report decoder. The proposed framework selectively emphasizes clinically significant image regions while suppressing redundant background information, thereby strengthening the correspondence between radiographic findings and generated textual descriptions. Experiments were conducted using chest radiograph–report pairs, with performance evaluated using standard natural-language-generation and clinical-consistency measures. The proposed AG-VLM achieved a BLEU-1 score of 0.521, BLEU-2 of 0.387, BLEU-3 of 0.301, BLEU-4 of 0.243, METEOR of 0.286, ROUGE-L of 0.418, and CIDEr of 0.472. For clinical content preservation, the framework obtained a clinical precision of 0.861, recall of 0.842, and F1-score of 0.851. The attention-guided architecture also achieved an abnormality localization accuracy of 91.7% and an overall clinical finding accuracy of 92.4%. Compared with the selected baseline vision-language report-generation model, AG-VLM improved BLEU-4 by 12.5%, METEOR by 9.6%, ROUGE-L by 8.3%, and clinical F1-score by 7.9%. These results indicate that explicit attention-guided visual reasoning combined with cross-modal semantic alignment can generate more accurate, clinically coherent, and contextually relevant radiology reports. The proposed framework therefore provides a scalable foundation for computer-assisted radiology reporting while retaining the need for radiologist verification before clinical use.