Uncertainty-Aware Cross-Modal Feature Interaction Fusion for Face and Full-Hand Biometric Identification
DOI:
https://doi.org/10.70917/ijcisim-2026-3749Keywords:
Multimodal Biometrics, Deep Learning, Feature Fusion, Attention Mechanism, Uncertainty Estimation, Convolutional Neural NetworksAbstract
Combining more than one biometric trait usually gives more reliable recognition than a single trait, but how the features are merged matters. Most systems simply concatenate the feature vectors, discarding information about how the two traits relate. We address this with UA-CFIF (Uncertainty-Aware Cross-Modal Feature Interaction Fusion), which fuses face and full hand images while explicitly modelling the interaction between them. Its Cross-Modal Interaction Module builds four views of the embeddings — the face, the hand, their difference, and their product — and a gated attention layer decides how much each should count. A dropout-based step then gives a sense of how sure the model is about each prediction. Where earlier MULBv1 studies cropped the hand to the palm, we keep the whole hand, fingers included. Across 528 test pairs from 176 people, UA-CFIF reaches 99.62% Rank-1 accuracy, 1.0000 AUC, and 0.15% EER, and unlike plain concatenation, it gives a clearer signal when it might be wrong.