r/Arabic_NLP 2d ago

👋 Welcome to r/Arabic_NLP !

1 Upvotes

Hey everyone!

This community is for anyone working on or interested in Arabic Natural Language Processing - researchers, engineers, students, and hobbyists alike.

What to post here:

  • Papers, preprints, and research summaries
  • Datasets, benchmarks, and evaluation results
  • Tools, libraries, and open-source projects
  • Questions about MSA, dialects, diacritization, morphology, MT, ASR, LLMs for Arabic, etc.
  • Job postings and collaboration requests
  • General discussion on the state of Arabic NLP

A few quick guidelines:

  • Keep posts on-topic and give them descriptive titles
  • Cite sources when sharing claims or results
  • Both Arabic and English are welcome
  • Check the sidebar for the full rule list before posting

Feel free to introduce yourself in the comments; what you work on, what dialect/domain interests you, or what brought you here. Looking forward to building this out with you.

Thanks for being part of the very first wave. Together, let's make r/Arabic_NLP amazing.


r/Arabic_NLP 14h ago

AraSentEval 2026: the overview paper + a look at the participating systems

1 Upvotes

The official overview paper is out: "AraSentEval 2026: A Shared Task on Sentiment Analysis and Swapping in Arabic" by Saad Ezzini, Shadi Abudalfa, Maram I. Alharbi, Salmane Chafik, Hamzah Luqman, Mo El-Haj, Paul Rayson, and Reem Alotaibi - published in the OSACT7 proceedings @ LREC 2026 (Palma, Mallorca, 11–16 May 2026). Full proceedings PDF: http://lrec-conf.org/proceedings/lrec2026/workshops/osact/2026.osact-1.0.pdf.

For context: the task had two subtasks: Arabic dialect sentiment classification on the new MDS-3 dataset (Moroccan/Egyptian/Jordanian/Saudi hotel reviews, 3-class), and sentiment swap on MA'AKS (rewriting a sentence to invert polarity while preserving meaning). It drew 19 final-phase participations total.

Here's every system paper that made it into the proceedings:

Subtask 1: Arabic Dialect Sentiment Classification

  • TTLab at AraSentEval: SARF (صرف) Sentiment Analysis via Root-based Fusion for Multi-Dialectal Arabic: Ali Abusaleh, Bhuvanesh Verma, Alexander Mehler: this one got picked as the workshop's lightning talk for the task
  • CasbAI at AraSentEval 2026: Robust Dialectal Arabic Sentiment Classification via Multi-Seed Ensembling and Data Augmentation: Chaima Abdelaziz, KahinaHouda Saadaoui, Faiza Belbachir, Lynda Said Lhadj
  • BDSI at AraSentEval Shared Task: A Multi-Transformer Contrastive Learning for Arabic Dialect Sentiment Analysis: Mohamed M'haouach, Kaouthar Elyoussoufi, Abdessamad Benlahbib, Hamza Alami
  • L3IA-Subtask 1 at AraSentEval Shared Task: Multi-Dialect Arabic Sentiment Classification via a Transformer-Based Approach: Mohamed M'haouach, Kaouthar Elyoussoufi, Hamza Alami, Abdessamad Benlahbib
  • University of Tripoli at AraSentEval: Fine-Tuning MARBERTv2 and CAMELBERT for Multi-Dialect Arabic Sentiment Analysis: Abdusalam F. Ahmad Nwesri, Amani Bahlul Sharif, Sarah Farag S. Hmeid
  • LinguArabic at AraSentEval 2026: MARBERT for Multi-Dialect Arabic Sentiment Analysis: Norah Saud Alshahrani, Elham Abdullah Al-Qarni, Shatha Hussan Alshomrani

Subtask 2: Arabic Sentiment Swap

  • A Comparative Study of Arabic Sentiment Swap Models for AraSentEval 2026: Yumna Hamdy, Mohab ElDamhougy, Yomna Eid, Ensaf Hussein
  • L3IA at AraSentEval 2026 Subtask 2: LLM-Based Multi-Step Pipeline for Arabic Sentiment Swap : Abdessamad Benlahbib, Hamza Alami, Mohamed M'haouach, Kaouthar Elyoussoufi
  • Codezone Research Group at AraSentEval Shared Task: Arabic Sentiment Swap beyond Negation Prepending, Benchmarking Multilingual T5 against Large Language Models on the MA'AKS Corpus: Abdulkadir Shehu Bichi, Sarah Yassine

A few patterns worth noting just from the titles: MARBERT/CAMELBERT fine-tuning and ensembling dominate the classification side, while the swap subtask splits between LLM pipelines and more classic seq2seq (T5) baselines: useful if you're weighing which approach to try on a similar dialectal task. For exact leaderboard rankings and metrics, check the overview paper linked above.

Sources:


r/Arabic_NLP 1d ago

Resource spotlight: Masader: the largest catalogue of Arabic NLP datasets

1 Upvotes

If you're not already using it, Masader is worth bookmarking. It's the largest public catalogue of Arabic NLP and speech datasets: over 1000+ datasets, each annotated with 25+ metadata fields (dialect, domain, source, license, volume, tasks, access type, paper link, citation count, and more).

What makes it genuinely useful over just googling around:

  • Filterable/searchable by dialect, domain, task, license, and access type: handy when you need something specific like "free, human-annotated, Levantine dialect, sentiment"
  • Each entry links the paper, the data host, and citation counts
  • Programmatic access via Hugging Face: datasets.load_dataset('arbml/masader')
  • Has a chat interface (Ask Masader) for querying the catalogue conversationally
  • Actively maintained: started in 2021 as part of the BigScience project, now maintained by the ARBML team and community, with a form to submit new datasets

Originally described in the paper "Masader: Metadata Sourcing for Arabic Text and Speech Data Resources" (Alyafeai et al., 2021): https://arxiv.org/abs/2110.06744, later expanded in "Masader Plus" (2022): https://arxiv.org/abs/2208.00932

GitHub (to contribute or browse the raw metadata): https://github.com/ARBML/masader

Sources:


r/Arabic_NLP 2d ago

SIGARAB Weekly Roundup (Jul 14–21): Emergency Reviewers Needed, Shared Task Updates

1 Upvotes

A few notable announcements from the SIGARAB mailing list this week: sharing here for anyone in the community who isn't subscribed.

Call for Emergency Reviewers: ArabicNLP 2026 (Samar Magdy, Jul 18)
The ArabicNLP 2026 Program Committee is short on reviewers and looking for qualified volunteers to do a quick-turnaround review of one or more papers.
Sign up: https://forms.gle/SsQg1WDtNFKQiEon7

KnowledgeGraphEval 2026 Shared Task: Webinar Recordings (Nagham Hamad, Jul 15)
Recordings from two webinars covering the subtasks, datasets, submission format, evaluation process, and Q&A are now up for anyone considering participating.

IslamicEval 2026: All Training and Dev Sets Are Ready (Tamer Elsayed, Jul 14)
Training and dev data for all subtasks are now available to registered teams via separate CodaBench competitions per subtask.

Call for Participation: HalluScoring 2026 @ ArabicNLP 2026 (Bouchekif Abdessalam, final call Jul 14)
Last call to join this shared task on hallucination detection and factuality verification for Arabic QA, covering Islamic Knowledge and General Culture domains across two tracks.
Site: https://halluscoring.github.io/HalluScoring-2026/

Sourced from the SIGARAB mailing list.