ArabicNLP 2026 Shared Task

ArA-DF 2026

Arabic speech deepfake detection across dialect and robustness conditions. Evaluation is complete, post-challenge gold labels are released, and system-description papers are next.

Task
Binary audio classification
Metric
Equal Error Rate
Venue
ArabicNLP 2026 @ EMNLP

Evaluation is complete and post-challenge resources are live.

  • Dataset live
  • Gold labels released
  • System papers due Aug 8
  • Camera-ready due Aug 22

Overview

A focused benchmark for Arabic audio authenticity.

ArA-DF 2026 is listed by ArabicNLP 2026 as Shared Task 14. Participating systems assign each Arabic speech sample to one of two labels: bona fide, for genuine human speech, or spoofed, for synthetic speech produced by text-to-speech or voice-conversion systems.

The task covers Modern Standard Arabic and Arabic dialects, with evaluation designed to reward detectors that generalize beyond narrow speakers, synthesis systems, and recording conditions.

Data & Baseline

Live dataset and baseline resources.

The ArA-DF 2026 dataset and baseline are hosted by ArabicSpeech on Hugging Face. Full loading commands and repository instructions remain on the Hugging Face cards.

Dataset on Hugging Face

Train/dev labels plus post-challenge gold labels for development-test and final-test configs.

Baseline on Hugging Face

wav2vec2 XLS-R 300M frontend with an AASIST backend and helper scripts.

Codabench leaderboards

Competition pages for the two released tracks; leaderboards and submission phases are maintained there.

ArabicNLP shared-tasks page

Conference-wide dates, contact listing, and paper-submission updates.

Dataset snapshot

Labels use 0 for spoofed or synthetic speech and 1 for bona fide speech. Released audio is 16 kHz mono PCM, packaged as lossless FLAC in WebDataset TAR shards, with metadata available through Hugging Face Parquet configs.

Dataset split rows and label availability
Split Rows Labels
Train22,500Yes
Dev21,000Yes
Track 1 development-test16,023Yes, in labeled config
Track 1 final-test144,210Yes, in labeled config
Track 2 development-test14,193Yes, in labeled config
Track 2 final-test127,746Yes, in labeled config

Baseline snapshot

The baseline repository includes a trained checkpoint, XLS-R frontend weights, configuration, parquet-generation utility, and published result file. Lower EER is better.

Baseline model-card EER results
Baseline split EER (%) Utterances
Track 1 development-test14.6716,023
Track 1 final-test14.54144,210
Track 2 development-test27.2114,193
Track 2 final-test27.11127,746

Values are baseline model-card results, not participant rankings.

Released August 4, 2026

Post-challenge test labels are now available.

The original challenge-era test configs remain blind for reproducibility. Gold labels are provided in separate Hugging Face configs whose names end in _labeled. The WebDataset TAR shards and JSON sidecars are unchanged and remain label-free.

Released gold-label counts
Labeled config Rows Label 0 Label 1
track-1_development_test_labeled16,02311,9264,097
track-1_test_labeled144,210107,30736,903
track-2_development_test_labeled14,19310,6833,510
track-2_test_labeled127,74696,25631,490

System Papers

System-description paper guidance.

Evaluation and leaderboard freeze are complete. Participating teams should now prepare the non-anonymous ArabicNLP system-description paper for their ArA-DF 2026 submission.

Due next

Initial system paper

, 11:59 PM UTC-12.

Venue status

Submission link pending

The ArabicNLP shared-tasks page currently lists the ArA-DF system-paper link as TBD.

Final version

Camera-ready deadline

, 11:59 PM UTC-12.

Required format

  • ArabicNLP/EMNLP short-paper format.
  • Maximum 4 pages, excluding unlimited references.
  • Non-anonymous submission.
  • Use the title format: <Team Name> at ArA-DF 2026: <Your Contribution>.
  • Use official ACL style files without modifying margins, fonts, or paper size.

What to include

  • System architecture, algorithms, and design decisions.
  • Data usage, preprocessing, training setup, tools, and versions.
  • Official submitted results and any clearly marked post-submission results.
  • Analysis, ablations, error patterns, limitations, and ethical considerations.
  • Code or model URL when available, to support reproducibility.

Recommended structure

  1. Abstract
  2. Introduction
  3. Background / Task Setup
  4. System Overview
  5. Experimental Setup
  6. Results and Analysis
  7. Conclusion
  8. Acknowledgments
  9. Appendix, if needed

Important Dates

Task timeline.

Completed challenge milestones and remaining system-paper deadlines.

Resources released

Complete Training, dev and open test data, evaluation scripts, and baseline.

Development phase

Complete Experimentation and submissions on the development-test sets.

Evaluation phase

Complete Submissions on the final-test sets and official scoring.

Leaderboard freeze

Complete Final leaderboard state is frozen for official results.

Gold labels released

New Post-challenge test labels released through Hugging Face labeled configs.

System papers due

Upcoming System description paper submissions due.

Acceptance notification

Upcoming Notification of acceptance for system-description papers.

Final results released

Upcoming Official final results are released to participants.

Camera-ready due

Upcoming Camera-ready versions of accepted system papers due.

All deadlines are 11:59 PM UTC-12 unless otherwise stated. The official ArA-DF system-paper submission link will be shared through the Google Group when confirmed.

Evaluation

Ranked by Equal Error Rate.

Equal Error Rate is the primary metric because it is threshold-independent and widely used in audio anti-spoofing evaluation. Lower EER is better; official scores came from the blind final-test submissions during the evaluation phase.

EER Primary ranking metric
Accuracy Secondary leaderboard metric
Macro-F1 Secondary leaderboard metric

Contact & Forum

Where to ask questions and report issues.

Use the public Google Group for task questions, discussions, announcements, and clarifications. Route private, platform, and repository issues to the appropriate channel below.

Registration

Team registration form

Register your team first, then request access on the Track 1 and/or Track 2 Codabench pages using the same team name and contact email.

Public forum

Questions, discussions, and announcements

Use the ArA-DF 2026 Google Group for participant-wide questions, official clarifications, system-paper questions, and shared announcements.

Private issues

Team-specific or sensitive support

Email the organizers for private issues such as team changes, registration mistakes, or sensitive access problems.

Submission platform

Codabench submissions and leaderboards

Use the corresponding Codabench competition page for track-specific submission, leaderboard, and platform issues.

Dataset and baseline

Hugging Face repository issues

Use the Hugging Face dataset and baseline repositories for data, loading, baseline, and model-card questions.

Organizers

Organizing team.

For public task questions, discussions, and announcements, use the ArA-DF 2026 Google Group: ara-df-2026@googlegroups.com.

Vasista Sai Lodagala

HUMAIN, Saudi Arabia

Yassine El Kheir

DFKI, Germany

Sara Althubaiti

HUMAIN, Saudi Arabia

Pedro Moreno Mengibar

HUMAIN, Saudi Arabia

Ahmed Ali

HUMAIN, Saudi Arabia