ORIGINAL ARTICLE
Building and Evaluating an Arabic Social-Media Dataset for Deepfake
Text Detection
More details
Hide details
1
Department of Information & Computer Science, King Fahd University of Petroleum & Minerals
Submission date: 2025-11-23
Final revision date: 2026-02-25
Acceptance date: 2026-05-03
Publication date: 2026-09-07
Corresponding author
Tarek Helmy
Department of Information & Computer Science, King Fahd University of Petroleum & Minerals
Journal of Undergraduate Research International 2026;2(2):48-55
KEYWORDS
TOPICS
ABSTRACT
Rapid advances in large language models have enabled the creation of sophisticated deepfake text, which poses a significant threat to social media information integrity. Arabic deepfake text detection is underexplored owing to the morphological complexity of the language and its diverse regional dialects. Studies in this domain often rely on narrow datasets or focus exclusively on Modern Standard Arabic and fail to capture the informal linguistic nuances prevalent on platforms like X. This study addresses these gaps by compiling a high-quality Arabic dataset spanning six major dialects and seven distinct domains, providing a representative benchmark for real-world environments. Real Arabic texts were collected from X, deepfake texts were generated via a prompt-based approach using GPT-4o-mini across multiple deception styles, and a pretrained BERT-base-uncased model was fine-tuned for binary classification using the Hugging Face Transformers framework. Experimental results validated the high quality of the proposed dataset, The model achieved training, validation, and test accuracies of 93.7%, 93.2%, and 92.5%, respectively, with an F1 score of 92.5%. A significant finding of this study was the variation in performance across different data sources. The highest accuracy was observed in entertainment content from YouTube, suggesting that the model effectively identifies the stylistic and expressive markers inherent in synthetic conversational Arabic. Our contributions include the curation of an authentic multi-dialectal corpus and validation of transformer-based architectures tailored for identifying AI-manipulated text. This study provides a benchmark for the Arabic natural language processing community.