Bridging Arabic NLP and Language Learning: A Critical Review of LLM-Based Approaches, Challenges, and Pedagogical Opportunities

Authors

  • Muh. Sabilar Rosyad Universitas Islam Negeri Sunan Kalijaga Yogyakarta, Indonesia Indonesia Flag Indonesia Author https://orcid.org/0000-0002-9733-743X
  • Laila Farah Fitria Universitas Islam Negeri Sunan Kalijaga Yogyakarta, Indonesia Indonesia Flag Indonesia Author

DOI:

https://doi.org/10.53515/ts5p7j52

Keywords:

Arabic Language Learning, Arabic Natural Language Processing, Computer-Assisted Language Learning (CALL), Large Language Models (LLMs), Second Language Acquisition (SLA)

Abstract

Recent advances in Arabic Natural Language Processing (NLP) have been largely driven by transformer-based architectures and Large Language Models (LLMs), which outperform traditional approaches across tasks such as sentiment analysis, machine translation, text classification, and question answering. Nevertheless, Arabic NLP research remains predominantly technocentric and insufficiently connected to Second Language Acquisition (SLA) theories and pedagogical frameworks, thereby limiting its educational applicability. This study aims to systematically synthesize recent developments in LLM-based Arabic NLP and critically examine their potential integration into Arabic language learning, particularly within SLA and Computer-Assisted Language Learning (CALL) frameworks. A Systematic Literature Review guided by the PRISMA framework was conducted using Scopus-indexed studies published between 2025 and 2026. The selected studies were analyzed through thematic synthesis across five dimensions: model architecture, task domain, data strategies, linguistic focus, and pedagogical relevance. The findings demonstrate the growing dominance of LLMs, including AraBERT, AraGPT2, and ALLAM, alongside increased reliance on data augmentation, synthetic corpus generation, and dialect-specific adaptation. Evaluation practices are also beginning to extend beyond conventional performance metrics toward trustworthiness, safety, reasoning, and robustness. However, the review identifies a near absence of SLA-informed applications, Arabic LLM-based CALL systems, and human-centered evaluation involving usability, learner engagement, trust, and cognitive load. Despite this gap, LLMs offer substantial pedagogical potential through adaptive feedback, contextualized input, interactive dialogue, and personalized learning support. The study concludes that Arabic NLP requires stronger interdisciplinary integration with SLA and CALL to transform technological capabilities into learner-centered, pedagogically grounded, and empirically validated language-learning applications.

References

Abdhood, S. F., Omar, N., & Tiun, S. (2025). A novel data augmentation framework for Arabic multi-label text classification using AraBART, AraGPT2, and Borderline-SMOTE. IEEE Access, 13, 169769–169778. https://doi.org/10.1109/ACCESS.2025.3609462

Aftan, S., Zhuang, Y., Aseeri, A. O., & Shah, H. (2026). A survey of natural language processing for classification of Saudi Arabic dialect: Advancements, opportunities, and challenges. Lecture Notes of the Institute for Computer Sciences, Social-Informatics and Telecommunications Engineering, 623, 105–124. https://doi.org/10.1007/978-3-031-92625-9_8

Ahmed, S., Allam, A., Hamdi, A., & Mohammed, A. (2025). Arabic symptom classification and diagnosis using transformer models and LLM-based augmentation. In 2025 3rd International Conference on Intelligent Methods, Systems, and Applications (IMSA) (pp. 18–23). IEEE. https://doi.org/10.1109/IMSA65733.2025.11167759

Al-Shaibani, M. S., & Ahmed, M. (2026). Arabic machine-generated text detection: Stylometric analysis and cross-model evaluation. Expert Systems with Applications, 305, Article 130644. https://doi.org/10.1016/j.eswa.2025.130644

Al-Thubaity, A. (2025). A novel dataset for Arabic domain-specific term extraction and comparative evaluation of BERT-based models for Arabic term extraction. ACM Transactions on Asian and Low-Resource Language Information Processing, 24(9). https://doi.org/10.1145/3748323

Aldawsari, M., & Dawood, O. (2025). AraEventCoref: An Arabic event coreference dataset and LLM benchmarks. ACM Transactions on Asian and Low-Resource Language Information Processing, 24(7). https://doi.org/10.1145/3743047

Alghamdi, E. A., Masoud, R. I., Alnuhait, D., Alomairi, A. Y., Ashraf, A., & Zaytoon, M. (2025). AraTrust: An evaluation of trustworthiness for LLMs in Arabic. In Proceedings of the International Conference on Computational Linguistics (pp. 8664–8679).

Almeman, K. (2025). Automated building of a multidialectal parallel Arabic corpus using large language models. Data, 10(12), Article 208. https://doi.org/10.3390/data10120208

Alrashidi, F., & Mathkour, H. I. (2026). An empirical study of transformer-based neural machine translation for English to Arabic. Information, 17(2), Article 198. https://doi.org/10.3390/info17020198

Alrayzah, A., Alsolami, F., & Saleh, M. (2026). AraFastQA: A transformer model for question answering for the Arabic language using few-shot learning. Computer Speech & Language, 95, Article 101857. https://doi.org/10.1016/j.csl.2025.101857

Alshahrani, E. S., & Aksoy, M. S. (2025). Adversarially robust multitask learning for offensive and hate speech detection in Arabic text using transformer-based models and RNN architectures. Applied Sciences, 15(17), Article 9602. https://doi.org/10.3390/app15179602

Alsolami, F., & Alrayzah, A. (2025). Arabic WikiTableQA: Benchmarking question answering over Arabic tables using large language models. Electronics, 14(19), Article 3829. https://doi.org/10.3390/electronics14193829

Bakr, M. A., Hany, M., Osama, A., & Gamal, N. (2025). A transformer-driven bilingual approach to Arabic text summarization. In 2025 3rd International Conference on Intelligent Methods, Systems, and Applications (IMSA) (pp. 112–117). IEEE. https://doi.org/10.1109/IMSA65733.2025.11167880

Boulesnam, I., & Boucetti, R. (2025). Arabic language characteristics that make its automatic processing challenging. International Arab Journal of Information Technology, 22(4), 814–831. https://doi.org/10.34028/iajit/22/4/14

Dahou, A., Dahou, A. H., Cheragui, M. A., Abdedaiem, A., Al-Qaness, M. A. A., Elaziz, M. A., Ewees, A. A., & Zheng, Z. (2025). A survey on dialect Arabic processing and analysis: Recent advances and future trends. ACM Transactions on Asian and Low-Resource Language Information Processing, 24(8). https://doi.org/10.1145/3747290

David, I., & Gelbard, R. (2025). Using machine learning for systematic literature review: Case in point, agile software development. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 15(1), Article e1569. https://doi.org/10.1002/widm.1569

Essameldin, O. A., Elbeih, A. O., Gomaa, W. H., & Elsersy, W. F. (2025). Arabic dialect classification using RNNs, transformers, and large language models: A comparative analysis. In 2025 3rd International Conference on Intelligent Methods, Systems, and Applications (IMSA) (pp. 472–477). IEEE. https://doi.org/10.1109/IMSA65733.2025.11167859

Ferroud, C., Maghfour, M., & Elouardighi, A. (2026). A comparative study of lexicon-based, machine learning, deep learning, and LLM methods for sentiment analysis on standard and dialectal Arabic texts. Lecture Notes in Networks and Systems, 1640, 23–35. https://doi.org/10.1007/978-3-032-07785-1_3

Galhom, A., Abobakr, A., Noseer, M., Abdalgwad, S., Nour, R., & Fares, A. (2025). AASTE: Arabic aspect sentiment triplet extraction. In 2025 7th Novel Intelligent and Leading Emerging Sciences Conference (NILES) (pp. 453–456). IEEE. https://doi.org/10.1109/NILES68063.2025.11232333

Guessoum, A., Berkani, L., Hadj Ameur, M. S., & Aouichat, A. (2026). Artificial intelligence and large language models for a new education and higher education paradigm in the Arab world. In Higher education in the Arab world: Artificial intelligence (pp. 373–397). Springer. https://doi.org/10.1007/978-3-031-99068-7_15

Hamed, I., Sabty, C., Abdennadher, S., Vu, N. T., Solorio, T., & Habash, N. (2025). A survey of code-switched Arabic NLP: Progress, challenges, and future directions. In Proceedings of the International Conference on Computational Linguistics (pp. 4561–4585).

Hasanaath, A., Alansari, A., Ashraf, A., Salmane, C., Luqman, H., & Ezzini, S. (2025). AraReasoner: Evaluating reasoning-based LLMs for Arabic NLP. In Findings of the Association for Computational Linguistics: EMNLP 2025 (pp. 18898–18914). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-emnlp.1028

Hithnawi, R. I., Hamarsheh, M. M. N., & Maree, M. (2025). AraBERT for Arabic cyberbullying detection in Facebook comments. Journal of Cybersecurity, 11(1), Article tyaf030. https://doi.org/10.1093/cybsec/tyaf030

Ibrahim, M., Gervás, P., & Méndez, G. (2026). AI-Cinema: A hybrid framework for Arabic movie scenario generation with traditional storytelling and cultural dialogs. Complexity, 2026(1), Article 9978799. https://doi.org/10.1155/cplx/9978799

Khader, K. A., Hussein, M. S., & Abu-Issa, A. S. (2025). Adapting large language models for Arabic: Comparative evaluation, fine-tuning, and ethical deployment. IEEE Access, 13, 182621–182632. https://doi.org/10.1109/ACCESS.2025.3623796

Kitchenham, B., Pearl Brereton, O., Budgen, D., Turner, M., Bailey, J., & Linkman, S. (2009). Systematic literature reviews in software engineering: A systematic literature review. Information and Software Technology, 51(1), 7–15. https://doi.org/10.1016/j.infsof.2008.09.009

Mohawesh, R., AlQarni, A. A., Alkhushayni, S. M., Daradkeh, T., & Bany Salameh, H. (2025). A new multilingual framework for fake reviews detection based on a large language model. The Journal of Supercomputing, 81(10). https://doi.org/10.1007/s11227-025-07636-6

Mubarak, H., Mohamed, A., & Hawasly, M. (2025). AraSafe: Benchmarking safety in Arabic LLMs. In Findings of the Association for Computational Linguistics: EMNLP 2025 (pp. 9976–9992). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-emnlp.529

Nashir, W. A., Mohsen, A. M., Al-Shargabi, A. A., Nour, M. K., & Al-Onazi, B. B. (2025). A complete, multi-layered Quranic treebank dataset with hybrid syntactic annotations for classical Arabic processing. Data in Brief, 62, Article 111940. https://doi.org/10.1016/j.dib.2025.111940

Ouali, S., El Garouani, S., & Chajia, M. (2025). Integrating artificial intelligence into the Arabic medical domain: A review of current progress, challenges, and future directions. In 2025 International Conference on Circuit, Systems, and Communication (ICCSC). IEEE. https://doi.org/10.1109/ICCSC66714.2025.11135224

Page, M. J., Moher, D., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., Shamseer, L., Tetzlaff, J. M., Akl, E. A., Brennan, S. E., Chou, R., Glanville, J., Grimshaw, J. M., Hróbjartsson, A., Lalu, M. M., Li, T., Loder, E. W., Mayo-Wilson, E., McDonald, S., … McKenzie, J. E. (2021). PRISMA 2020 explanation and elaboration: Updated guidance and exemplars for reporting systematic reviews. BMJ, 372, Article n160. https://doi.org/10.1136/bmj.n160

Saiful Bari, M., Alnumay, Y., Alzahrani, N. A., Alotaibi, N. M., Alyahya, H. A., AlRashed, S., Mirza, F. A., Alsubaie, S. Z., Alahmed, H. A., Alabduljabbar, G., Alkhathran, R., Almushayqih, Y., Alnajim, R., Alsubaihi, S., Al Mansour, M., Alrubaian, M., Alammari, A., Alawami, Z., Al-Thubaity, A., … Khan, H. (2025). ALLAM: Large language models for Arabic and English. In The 13th International Conference on Learning Representations (ICLR 2025) (pp. 59235–59270).

Shang, G., Abdine, H., Khoubrane, Y., Mohamed, A., Abbahaddou, Y., Ennadir, S., Momayiz, I., Ren, X., Moulines, E., Nakov, P., Vazirgiannis, M., & Xing, E. (2025). Atlas-Chat: Adapting large language models for low-resource Moroccan Arabic dialect. In Proceedings of the International Conference on Computational Linguistics (pp. 9–30).

Shi, Z., & Agrawal, R. (2025). A comprehensive survey of contemporary Arabic sentiment analysis: Methods, challenges, and future directions. In Findings of the Association for Computational Linguistics: NAACL 2025 (pp. 3760–3772). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-naacl.208

Snyder, H. (2019). Literature review as a research methodology: An overview and guidelines. Journal of Business Research, 104, 333–339. https://doi.org/10.1016/j.jbusres.2019.07.039

Tahtah, O., Zinedine, A., & Fardousse, K. (2025). Arabic legal text classification using pre-trained transformers and deep learning. In 2025 International Conference on Circuit, Systems, and Communication (ICCSC). IEEE. https://doi.org/10.1109/ICCSC66714.2025.11135374

Thomas, J., & Harden, A. (2008). Methods for the thematic synthesis of qualitative research in systematic reviews. BMC Medical Research Methodology, 8, Article 45. https://doi.org/10.1186/1471-2288-8-45

Downloads

Published

2026-06-10

How to Cite

Bridging Arabic NLP and Language Learning: A Critical Review of LLM-Based Approaches, Challenges, and Pedagogical Opportunities. (2026). Learning, Media and Technology in Arabic Education, 2(1), 1-12. https://doi.org/10.53515/ts5p7j52