Dataset on Hugging Face
Train/dev labels plus post-challenge gold labels for development-test and final-test configs.
ArabicNLP 2026 Shared Task
Arabic speech deepfake detection across dialect and robustness conditions. Evaluation is complete, post-challenge gold labels are released, and system-description papers are next.
Evaluation is complete and post-challenge resources are live.
Overview
ArA-DF 2026 is listed by ArabicNLP 2026 as Shared Task 14. Participating systems assign each Arabic speech sample to one of two labels: bona fide, for genuine human speech, or spoofed, for synthetic speech produced by text-to-speech or voice-conversion systems.
The task covers Modern Standard Arabic and Arabic dialects, with evaluation designed to reward detectors that generalize beyond narrow speakers, synthesis systems, and recording conditions.
Tracks
Both tracks use Equal Error Rate as the primary ranking metric, with Accuracy and macro-F1 reported for additional interpretation.
Evaluate whether systems remain reliable on Arabic dialects and speakers that are not represented in the training data.
Open Codabench Track 2Test robustness under practical audio conditions such as compression artifacts, background noise, and re-recording effects.
Open CodabenchData & Baseline
The ArA-DF 2026 dataset and baseline are hosted by ArabicSpeech on Hugging Face. Full loading commands and repository instructions remain on the Hugging Face cards.
Train/dev labels plus post-challenge gold labels for development-test and final-test configs.
wav2vec2 XLS-R 300M frontend with an AASIST backend and helper scripts.
Competition pages for the two released tracks; leaderboards and submission phases are maintained there.
Conference-wide dates, contact listing, and paper-submission updates.
Labels use 0 for spoofed or synthetic speech and 1 for bona fide speech. Released audio is 16 kHz mono PCM, packaged as lossless FLAC in WebDataset TAR shards, with metadata available through Hugging Face Parquet configs.
| Split | Rows | Labels |
|---|---|---|
| Train | 22,500 | Yes |
| Dev | 21,000 | Yes |
| Track 1 development-test | 16,023 | Yes, in labeled config |
| Track 1 final-test | 144,210 | Yes, in labeled config |
| Track 2 development-test | 14,193 | Yes, in labeled config |
| Track 2 final-test | 127,746 | Yes, in labeled config |
The baseline repository includes a trained checkpoint, XLS-R frontend weights, configuration, parquet-generation utility, and published result file. Lower EER is better.
| Baseline split | EER (%) | Utterances |
|---|---|---|
| Track 1 development-test | 14.67 | 16,023 |
| Track 1 final-test | 14.54 | 144,210 |
| Track 2 development-test | 27.21 | 14,193 |
| Track 2 final-test | 27.11 | 127,746 |
Values are baseline model-card results, not participant rankings.
Released August 4, 2026
The original challenge-era test configs remain blind for reproducibility. Gold labels
are provided in separate Hugging Face configs whose names end in _labeled.
The WebDataset TAR shards and JSON sidecars are unchanged and remain label-free.
| Labeled config | Rows | Label 0 | Label 1 |
|---|---|---|---|
track-1_development_test_labeled | 16,023 | 11,926 | 4,097 |
track-1_test_labeled | 144,210 | 107,307 | 36,903 |
track-2_development_test_labeled | 14,193 | 10,683 | 3,510 |
track-2_test_labeled | 127,746 | 96,256 | 31,490 |
System Papers
Evaluation and leaderboard freeze are complete. Participating teams should now prepare the non-anonymous ArabicNLP system-description paper for their ArA-DF 2026 submission.
, 11:59 PM UTC-12.
The ArabicNLP shared-tasks page currently lists the ArA-DF system-paper link as TBD.
, 11:59 PM UTC-12.
<Team Name> at ArA-DF 2026: <Your Contribution>.Teams need OpenReview accounts. Accepted system papers must be presented at ArabicNLP 2026, in person or virtually, to appear in the proceedings.
Important Dates
Completed challenge milestones and remaining system-paper deadlines.
Complete Training, dev and open test data, evaluation scripts, and baseline.
Complete Experimentation and submissions on the development-test sets.
Complete Submissions on the final-test sets and official scoring.
Complete Final leaderboard state is frozen for official results.
New Post-challenge test labels released through Hugging Face labeled configs.
Upcoming System description paper submissions due.
Upcoming Notification of acceptance for system-description papers.
Upcoming Official final results are released to participants.
Upcoming Camera-ready versions of accepted system papers due.
All deadlines are 11:59 PM UTC-12 unless otherwise stated. The official ArA-DF system-paper submission link will be shared through the Google Group when confirmed.
Evaluation
Equal Error Rate is the primary metric because it is threshold-independent and widely used in audio anti-spoofing evaluation. Lower EER is better; official scores came from the blind final-test submissions during the evaluation phase.
Contact & Forum
Use the public Google Group for task questions, discussions, announcements, and clarifications. Route private, platform, and repository issues to the appropriate channel below.
Register your team first, then request access on the Track 1 and/or Track 2 Codabench pages using the same team name and contact email.
Use the ArA-DF 2026 Google Group for participant-wide questions, official clarifications, system-paper questions, and shared announcements.
Email the organizers for private issues such as team changes, registration mistakes, or sensitive access problems.
Use the corresponding Codabench competition page for track-specific submission, leaderboard, and platform issues.
Use the Hugging Face dataset and baseline repositories for data, loading, baseline, and model-card questions.
Organizers
For public task questions, discussions, and announcements, use the ArA-DF 2026 Google Group: ara-df-2026@googlegroups.com.
HUMAIN, Saudi Arabia
DFKI, Germany
HUMAIN, Saudi Arabia
HUMAIN, Saudi Arabia
HUMAIN, Saudi Arabia