Spatio-Temporal Dual-Attention Fusion Network (STA-Net): A Novel Deep Learning Framework for Robust Deepfake Video Deception Detection

Authors

  • Kavita Lal Department of Computer Science and Engineering, [Poornima University], Jaipur, Rajasthan, India.
  • Savita Shiwani Department of Computer Science and Engineering, [Poornima University], Jaipur, Rajasthan, India.
  • Geeta Chhabra Gandhi Department of Computer Applications, [Manipal University, Jaipur], Jaipur, Rajasthan, India.

DOI:

https://doi.org/10.70917/ijcisim-2026-2671

Keywords:

Deepfake detection, spatio-temporal attention, cross-attention fusion, convolutional LSTM, video forensics, deep learning, media forensics, FaceForensics++

Abstract

The proliferation of generative adversarial networks (GANs) and diffusion-based synthesis tools has made hyper-realistic deepfake videos trivially accessible, posing a severe threat to information integrity, biometric security, and public trust in digital media. Existing detectors that rely purely on frame-level spatial artefacts or purely on inter-frame temporal cues generalize poorly to unseen manipulation types and compression levels. This paper proposes the Spatio-Temporal dual-Attention fusion Network (STA-Net), a novel end-to-end deep learning framework that jointly models fine-grained spatial forgery artefacts and long-range temporal inconsistencies through a learnable cross-attention fusion module. The spatial branch employs a convolutional encoder augmented with a spatial attention module (SAM) to localize blending boundaries, texture discontinuities, and warping artefacts within individual frames, while the temporal branch employs a bidirectional convolutional LSTM strengthened by a temporal attention module (TAM) to capture micro-expression irregularities, eye-blink asynchrony, and unnatural motion dynamics across frames. A cross-attention fusion block adaptively re-weights spatial and temporal embeddings before classification, allowing the network to emphasize whichever cue is more discriminative for a given manipulation type. STA-Net is evaluated on three widely used benchmarks — FaceForensics++, Celeb-DF (v2), and the Deepfake Detection Challenge (DFDC) preview dataset — and achieves 98.4% accuracy and an AUC of 0.988 on FaceForensics++, outperforming recent CNN-LSTM, two-stream, vision-transformer, and multi-view spatiotemporal transformer baselines by 1.1–4.6 percentage points while maintaining competitive inference latency (41 ms per 16-frame clip). Ablation studies confirm that the spatial and temporal attention modules are complementary, and that the cross-attention fusion strategy improves cross-dataset generalization by 3.2% over simple feature concatenation. The results establish STA-Net as an accurate, interpretable, and computationally practical solution for real-world deepfake video deception detection.

Downloads

Download data is not yet available.

Downloads

Published

2026-07-02

How to Cite

Kavita Lal, Savita Shiwani, & Geeta Chhabra Gandhi. (2026). Spatio-Temporal Dual-Attention Fusion Network (STA-Net): A Novel Deep Learning Framework for Robust Deepfake Video Deception Detection. International Journal of Computer Information Systems and Industrial Management Applications, 18(4s), 1410–1420. https://doi.org/10.70917/ijcisim-2026-2671

Issue

Section

Original Articles