Bridging Arabic NLP and Language Learning: A Critical Review of LLM-Based Approaches, Challenges, and Pedagogical Opportunities
DOI:
https://doi.org/10.53515/ts5p7j52Keywords:
Arabic Language Learning, Arabic Natural Language Processing, Computer-Assisted Language Learning (CALL), Large Language Models (LLMs), Second Language Acquisition (SLA)Abstract
Recent advances in Arabic Natural Language Processing (NLP) have been largely driven by transformer-based architectures and Large Language Models (LLMs), which outperform traditional approaches across tasks such as sentiment analysis, machine translation, text classification, and question answering. Nevertheless, Arabic NLP research remains predominantly technocentric and insufficiently connected to Second Language Acquisition (SLA) theories and pedagogical frameworks, thereby limiting its educational applicability. This study aims to systematically synthesize recent developments in LLM-based Arabic NLP and critically examine their potential integration into Arabic language learning, particularly within SLA and Computer-Assisted Language Learning (CALL) frameworks. A Systematic Literature Review guided by the PRISMA framework was conducted using Scopus-indexed studies published between 2025 and 2026. The selected studies were analyzed through thematic synthesis across five dimensions: model architecture, task domain, data strategies, linguistic focus, and pedagogical relevance. The findings demonstrate the growing dominance of LLMs, including AraBERT, AraGPT2, and ALLAM, alongside increased reliance on data augmentation, synthetic corpus generation, and dialect-specific adaptation. Evaluation practices are also beginning to extend beyond conventional performance metrics toward trustworthiness, safety, reasoning, and robustness. However, the review identifies a near absence of SLA-informed applications, Arabic LLM-based CALL systems, and human-centered evaluation involving usability, learner engagement, trust, and cognitive load. Despite this gap, LLMs offer substantial pedagogical potential through adaptive feedback, contextualized input, interactive dialogue, and personalized learning support. The study concludes that Arabic NLP requires stronger interdisciplinary integration with SLA and CALL to transform technological capabilities into learner-centered, pedagogically grounded, and empirically validated language-learning applications.
References
Abdhood, S. F., Omar, N., & Tiun, S. (2025). A novel data augmentation framework for Arabic multi-label text classification using AraBART, AraGPT2, and Borderline-SMOTE. IEEE Access, 13, 169769–169778. https://doi.org/10.1109/ACCESS.2025.3609462
Aftan, S., Zhuang, Y., Aseeri, A. O., & Shah, H. (2026). A survey of natural language processing for classification of Saudi Arabic dialect: Advancements, opportunities, and challenges. Lecture Notes of the Institute for Computer Sciences, Social-Informatics and Telecommunications Engineering, 623, 105–124. https://doi.org/10.1007/978-3-031-92625-9_8
Ahmed, S., Allam, A., Hamdi, A., & Mohammed, A. (2025). Arabic symptom classification and diagnosis using transformer models and LLM-based augmentation. In 2025 3rd International Conference on Intelligent Methods, Systems, and Applications (IMSA) (pp. 18–23). IEEE. https://doi.org/10.1109/IMSA65733.2025.11167759
Al-Shaibani, M. S., & Ahmed, M. (2026). Arabic machine-generated text detection: Stylometric analysis and cross-model evaluation. Expert Systems with Applications, 305, Article 130644. https://doi.org/10.1016/j.eswa.2025.130644
Al-Thubaity, A. (2025). A novel dataset for Arabic domain-specific term extraction and comparative evaluation of BERT-based models for Arabic term extraction. ACM Transactions on Asian and Low-Resource Language Information Processing, 24(9). https://doi.org/10.1145/3748323
Aldawsari, M., & Dawood, O. (2025). AraEventCoref: An Arabic event coreference dataset and LLM benchmarks. ACM Transactions on Asian and Low-Resource Language Information Processing, 24(7). https://doi.org/10.1145/3743047
Alghamdi, E. A., Masoud, R. I., Alnuhait, D., Alomairi, A. Y., Ashraf, A., & Zaytoon, M. (2025). AraTrust: An evaluation of trustworthiness for LLMs in Arabic. In Proceedings of the International Conference on Computational Linguistics (pp. 8664–8679).
Almeman, K. (2025). Automated building of a multidialectal parallel Arabic corpus using large language models. Data, 10(12), Article 208. https://doi.org/10.3390/data10120208
Alrashidi, F., & Mathkour, H. I. (2026). An empirical study of transformer-based neural machine translation for English to Arabic. Information, 17(2), Article 198. https://doi.org/10.3390/info17020198
Alrayzah, A., Alsolami, F., & Saleh, M. (2026). AraFastQA: A transformer model for question answering for the Arabic language using few-shot learning. Computer Speech & Language, 95, Article 101857. https://doi.org/10.1016/j.csl.2025.101857
Alshahrani, E. S., & Aksoy, M. S. (2025). Adversarially robust multitask learning for offensive and hate speech detection in Arabic text using transformer-based models and RNN architectures. Applied Sciences, 15(17), Article 9602. https://doi.org/10.3390/app15179602
Alsolami, F., & Alrayzah, A. (2025). Arabic WikiTableQA: Benchmarking question answering over Arabic tables using large language models. Electronics, 14(19), Article 3829. https://doi.org/10.3390/electronics14193829
Bakr, M. A., Hany, M., Osama, A., & Gamal, N. (2025). A transformer-driven bilingual approach to Arabic text summarization. In 2025 3rd International Conference on Intelligent Methods, Systems, and Applications (IMSA) (pp. 112–117). IEEE. https://doi.org/10.1109/IMSA65733.2025.11167880
Boulesnam, I., & Boucetti, R. (2025). Arabic language characteristics that make its automatic processing challenging. International Arab Journal of Information Technology, 22(4), 814–831. https://doi.org/10.34028/iajit/22/4/14
Dahou, A., Dahou, A. H., Cheragui, M. A., Abdedaiem, A., Al-Qaness, M. A. A., Elaziz, M. A., Ewees, A. A., & Zheng, Z. (2025). A survey on dialect Arabic processing and analysis: Recent advances and future trends. ACM Transactions on Asian and Low-Resource Language Information Processing, 24(8). https://doi.org/10.1145/3747290
David, I., & Gelbard, R. (2025). Using machine learning for systematic literature review: Case in point, agile software development. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 15(1), Article e1569. https://doi.org/10.1002/widm.1569
Essameldin, O. A., Elbeih, A. O., Gomaa, W. H., & Elsersy, W. F. (2025). Arabic dialect classification using RNNs, transformers, and large language models: A comparative analysis. In 2025 3rd International Conference on Intelligent Methods, Systems, and Applications (IMSA) (pp. 472–477). IEEE. https://doi.org/10.1109/IMSA65733.2025.11167859
Ferroud, C., Maghfour, M., & Elouardighi, A. (2026). A comparative study of lexicon-based, machine learning, deep learning, and LLM methods for sentiment analysis on standard and dialectal Arabic texts. Lecture Notes in Networks and Systems, 1640, 23–35. https://doi.org/10.1007/978-3-032-07785-1_3
Galhom, A., Abobakr, A., Noseer, M., Abdalgwad, S., Nour, R., & Fares, A. (2025). AASTE: Arabic aspect sentiment triplet extraction. In 2025 7th Novel Intelligent and Leading Emerging Sciences Conference (NILES) (pp. 453–456). IEEE. https://doi.org/10.1109/NILES68063.2025.11232333
Guessoum, A., Berkani, L., Hadj Ameur, M. S., & Aouichat, A. (2026). Artificial intelligence and large language models for a new education and higher education paradigm in the Arab world. In Higher education in the Arab world: Artificial intelligence (pp. 373–397). Springer. https://doi.org/10.1007/978-3-031-99068-7_15
Hamed, I., Sabty, C., Abdennadher, S., Vu, N. T., Solorio, T., & Habash, N. (2025). A survey of code-switched Arabic NLP: Progress, challenges, and future directions. In Proceedings of the International Conference on Computational Linguistics (pp. 4561–4585).
Hasanaath, A., Alansari, A., Ashraf, A., Salmane, C., Luqman, H., & Ezzini, S. (2025). AraReasoner: Evaluating reasoning-based LLMs for Arabic NLP. In Findings of the Association for Computational Linguistics: EMNLP 2025 (pp. 18898–18914). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-emnlp.1028
Hithnawi, R. I., Hamarsheh, M. M. N., & Maree, M. (2025). AraBERT for Arabic cyberbullying detection in Facebook comments. Journal of Cybersecurity, 11(1), Article tyaf030. https://doi.org/10.1093/cybsec/tyaf030
Ibrahim, M., Gervás, P., & Méndez, G. (2026). AI-Cinema: A hybrid framework for Arabic movie scenario generation with traditional storytelling and cultural dialogs. Complexity, 2026(1), Article 9978799. https://doi.org/10.1155/cplx/9978799
Khader, K. A., Hussein, M. S., & Abu-Issa, A. S. (2025). Adapting large language models for Arabic: Comparative evaluation, fine-tuning, and ethical deployment. IEEE Access, 13, 182621–182632. https://doi.org/10.1109/ACCESS.2025.3623796
Kitchenham, B., Pearl Brereton, O., Budgen, D., Turner, M., Bailey, J., & Linkman, S. (2009). Systematic literature reviews in software engineering: A systematic literature review. Information and Software Technology, 51(1), 7–15. https://doi.org/10.1016/j.infsof.2008.09.009
Mohawesh, R., AlQarni, A. A., Alkhushayni, S. M., Daradkeh, T., & Bany Salameh, H. (2025). A new multilingual framework for fake reviews detection based on a large language model. The Journal of Supercomputing, 81(10). https://doi.org/10.1007/s11227-025-07636-6
Mubarak, H., Mohamed, A., & Hawasly, M. (2025). AraSafe: Benchmarking safety in Arabic LLMs. In Findings of the Association for Computational Linguistics: EMNLP 2025 (pp. 9976–9992). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-emnlp.529
Nashir, W. A., Mohsen, A. M., Al-Shargabi, A. A., Nour, M. K., & Al-Onazi, B. B. (2025). A complete, multi-layered Quranic treebank dataset with hybrid syntactic annotations for classical Arabic processing. Data in Brief, 62, Article 111940. https://doi.org/10.1016/j.dib.2025.111940
Ouali, S., El Garouani, S., & Chajia, M. (2025). Integrating artificial intelligence into the Arabic medical domain: A review of current progress, challenges, and future directions. In 2025 International Conference on Circuit, Systems, and Communication (ICCSC). IEEE. https://doi.org/10.1109/ICCSC66714.2025.11135224
Page, M. J., Moher, D., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., Shamseer, L., Tetzlaff, J. M., Akl, E. A., Brennan, S. E., Chou, R., Glanville, J., Grimshaw, J. M., Hróbjartsson, A., Lalu, M. M., Li, T., Loder, E. W., Mayo-Wilson, E., McDonald, S., … McKenzie, J. E. (2021). PRISMA 2020 explanation and elaboration: Updated guidance and exemplars for reporting systematic reviews. BMJ, 372, Article n160. https://doi.org/10.1136/bmj.n160
Saiful Bari, M., Alnumay, Y., Alzahrani, N. A., Alotaibi, N. M., Alyahya, H. A., AlRashed, S., Mirza, F. A., Alsubaie, S. Z., Alahmed, H. A., Alabduljabbar, G., Alkhathran, R., Almushayqih, Y., Alnajim, R., Alsubaihi, S., Al Mansour, M., Alrubaian, M., Alammari, A., Alawami, Z., Al-Thubaity, A., … Khan, H. (2025). ALLAM: Large language models for Arabic and English. In The 13th International Conference on Learning Representations (ICLR 2025) (pp. 59235–59270).
Shang, G., Abdine, H., Khoubrane, Y., Mohamed, A., Abbahaddou, Y., Ennadir, S., Momayiz, I., Ren, X., Moulines, E., Nakov, P., Vazirgiannis, M., & Xing, E. (2025). Atlas-Chat: Adapting large language models for low-resource Moroccan Arabic dialect. In Proceedings of the International Conference on Computational Linguistics (pp. 9–30).
Shi, Z., & Agrawal, R. (2025). A comprehensive survey of contemporary Arabic sentiment analysis: Methods, challenges, and future directions. In Findings of the Association for Computational Linguistics: NAACL 2025 (pp. 3760–3772). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-naacl.208
Snyder, H. (2019). Literature review as a research methodology: An overview and guidelines. Journal of Business Research, 104, 333–339. https://doi.org/10.1016/j.jbusres.2019.07.039
Tahtah, O., Zinedine, A., & Fardousse, K. (2025). Arabic legal text classification using pre-trained transformers and deep learning. In 2025 International Conference on Circuit, Systems, and Communication (ICCSC). IEEE. https://doi.org/10.1109/ICCSC66714.2025.11135374
Thomas, J., & Harden, A. (2008). Methods for the thematic synthesis of qualitative research in systematic reviews. BMC Medical Research Methodology, 8, Article 45. https://doi.org/10.1186/1471-2288-8-45
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Muh. Sabilar Rosyad, Laila Farah Fitria (Author)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

All publications by the UNIKHAMS [e-ISSN: