Beyond Single-Corpus Leaderboards: Controlled Cross-Corpus Evaluation of Frequency-Gated and Multi-Scale Cross-Attention Networks for Image Deepfake Detection

Authors

  • Abhay Mundra Research Scholar, India.
  • Dr.Mamta Meena Research Scholar, India.

DOI:

https://doi.org/10.70917/ijcisim-2026-5393

Keywords:

Deepfake detection, media forensics, bidirectional LSTM, vision transformer, cross-attention, frequency-domain analysis, cross-dataset generalisation, Matthews correlation coefficient

Abstract

Image deepfake detectors are routinely selected on single-corpus accuracy, yet deployment requires a decision on media whose manipulation family is unknown. This paper asks whether single-corpus ranking predicts ranking under unknown provenance, and answers in the negative. Six architectures are trained from a common initialisation under one identical data, preprocessing, augmentation and optimisation protocol: four re-implement the established hybrid convolutional–recurrent family, and two are proposed here. FA-ViT-BiLSTMcouples vision-transformer patch tokens with log-magnitude Fourier tokens through a learned element-wise modality gate, allowing the spectral stream to be attenuated under compression. MSCA-CTNet exchanges information between three-scale convolutional tokens and global transformer tokens through bidirectional multi-head cross-attention. All six are evaluated on Celeb-DF v2, DFDC, UADFV, FaceForensics++ C23 and DeepDetect-2025, on a mixed-provenance Combined set, and under a leave-one-corpus-out protocol. Dual-backbone hybrids give the strongest single-corpus results, reaching 0.953 accuracy on DFDC, while MSCA-CTNet attains the highest ROC-AUC at 0.987. The central finding is a dissociation: the architecture that wins four of five corpora loses 0.203 accuracy under mixed provenance and ranks fourth on held-out corpora, where the two proposed architectures lead with mean AUCs of 0.849 and 0.830. Class-wise accuracy and the Matthews correlation coefficient are shown to be necessary rather than optional.

Downloads

Download data is not yet available.

Downloads

Published

2026-09-01

How to Cite

Abhay Mundra, & Dr.Mamta Meena. (2026). Beyond Single-Corpus Leaderboards: Controlled Cross-Corpus Evaluation of Frequency-Gated and Multi-Scale Cross-Attention Networks for Image Deepfake Detection. International Journal of Computer Information Systems and Industrial Management Applications, 18(21s), 1282–1316. https://doi.org/10.70917/ijcisim-2026-5393

Issue

Section

Original Articles