FEMT-GAT: An Explainable Federated Multimodal Transformer–Graph Attention Network for Privacy-Preserving Disease Prediction
DOI:
https://doi.org/10.70917/ijcisim-2026-3898Keywords:
Federated learning, multimodal learning, transformer, graph attention network, explainable artificial intelligence, differential privacy, secure aggregation, disease predictionAbstract
The growing use of multimodal clinical data creates an opportunity to improve disease prediction, but conventional centralized learning requires hospitals to pool sensitive patient records and often provides limited interpretability. This paper presents FEMT-GAT, an explainable Federated Multimodal Transformer–Graph Attention Network for privacy-preserving disease prediction across geographically distributed healthcare institutions. The framework uses modality-specific encoders for electronic health records, laboratory measurements, demographic variables, and medical images. A cross-modal transformer learns context-dependent interactions among available modalities, while a patient similarity graph and multi-head graph attention network incorporate relational evidence from clinically comparable cases. Collaborative training is performed through federated optimization so that raw data remain at the originating institutions. Update clipping, differential privacy, encryption, and secure aggregation are incorporated to reduce information leakage. The inference module produces disease probabilities together with modality-level, feature-level, image-level, and graph-level explanations using attention rollout, modality ablation, SHAP or integrated gradients, Grad-CAM, and graph attention coefficients. The proposed architecture supports incomplete modalities through masking and modality dropout and can be adapted to binary diagnosis, multiclass classification, risk regression, and survival analysis. FEMT-GAT therefore provides a unified methodology for accurate, transparent, and privacy-aware collaborative clinical intelligence.