XStack-Net: A stacking-based deep learning framework for robust deepfake detection


Ayata F., Tuna R. İ.

Cluster Computing, cilt.29, sa.8, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 29 Sayı: 8
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1007/s10586-026-06319-y
  • Dergi Adı: Cluster Computing
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Aerospace Database, Applied Science & Technology Source, Compendex, INSPEC, Academic Search Ultimate (EBSCO), Technology Collection (ProQuest)
  • Anahtar Kelimeler: CNN, Cross-dataset generalization, Deepfake detection, Frame-based deepfake analysis, Meta-learning, Stacking framework, XGBoost
  • Van Yüzüncü Yıl Üniversitesi Adresli: Evet

Özet

The rapid spread of deepfake videos raises significant concerns about digital security, media accuracy, and the reliability of information. To address this issue, this study develops a stacking-based deepfake detection framework called XStack-Net. The proposed framework combines three complementary convolutional neural network architectures—ResNet-50, DenseNet-121, and InceptionV3—and uses XGBoost as the meta-learner. Instead of directly processing full videos, the proposed architectural framework extracts five representative frames from each video and applies Dlib-based face detection and cropping operations to transform video data into a more manageable image dataset. Within this approach, a limited number of frames are selected from each video to extract face regions, creating a balanced image dataset for model training. This approach makes the data preparation and model training process more practical, faster, and computationally more efficient compared to direct video-level processing, while also allowing the model to focus on face regions associated with manipulation. The proposed framework is evaluated using a comprehensive experimental protocol including base model comparisons, meta-learner ablation analysis, analysis of variance, calibration evaluation, and computational cost analysis. In addition to LR, SVM, and ElasticNet, XGBoost was also examined as a final meta-learner, and the results showed that XGBoost’s nonlinear modeling capabilities enabled the most effective combination of basic model outputs. The proposed XGBoost-based stacking framework exhibited stronger classification performance compared to single deep learning models, achieving 98.25% accuracy and 99.81% AUC. Additional experiments conducted on Celeb-DF and OpenForensics datasets revealed strong results under the adopted evaluation protocol. However, performance degradation was observed under the domain shift effect in direct cross-dataset evaluations, indicating that generalization among heterogeneous data distributions remains a challenging problem. Overall, the results demonstrate that XStack-Net offers a practical, transparent, and empirically enhanced framework for frame-based deepfake detection.