Hybrid Framework for Pre-Assessing Arabic Multiple-Choice Questions Using ChatGPT and Traditional Item Analysis

Authors

DOI:

https://doi.org/10.23851/mjs.v37i3.1857

Keywords:

MCQ, Evaluation, ChatGPT, Psychometric analysis, Arabic assessment, AI assessment

Abstract

Background: Multiple-choice questions (MCQs) are widely used in educational evaluation; nonetheless, their effectiveness depends widely on object quality and psychometric properties. Traditional item analysis methods are typically applied after exam administration, reducing the ability to detect problematic items during the test development phase. Objective: This research suggests a hybrid framework that combines AI-based pre-assessment with traditional post-test psychometric analysis to validate Arabic MCQs before and after exam administration. Methods: The suggested framework uses ChatGPT-4 to conduct a pre-assessment of MCQ items by estimating item difficulty, discrimination potential, and distractor effectiveness based on linguistic and semantic properties. Traditional psychometric analysis was then performed using actual learner response data collected from 195 individuals to compute item difficulty, discrimination indices, reliability estimates, and distractor effectiveness. Results: A moderate positive correlation was observed between AI-based predictions and traditional psychometric results (r = 0.42, p<0.05). The AI model demonstrates effectiveness in identifying possible weaknesses in question form and distractor quality. Yet, its performance was limited in estimating discrimination indicators that depend directly on actual learner response behavior. Conclusions: The findings suggest that AI-assisted pre-assessment can support the early identification of problematic MCQ items before exam administration. Integrating AI-based linguistic evaluation with traditional psychometric analysis may enhance the quality of Arabic educational assessment.

Downloads

Download data is not yet available.

References

E. Emekli and B. N. Karahan, "AI in radiography education: Evaluating multiple-choice questions difficulty and discrimination," Journal of Medical Imaging and Radiation Sciences, vol. 56, no. 4, Art no. 101896, 2025.

Q. Zhang, Z. Huang, Y. Huang, G. Wang, R. Zhang, J. Yang, Y. Cheng, B. Chen, H. Wang, K. Qiu, et al., “Generative AI in medical education: Feasibility and educational value of LLM-generated clinical cases with MCQs,” BMC Medical Education, vol. 25, no. 1, Art no. 1502, 2025

C. Grévisse, M. A. S. Pavlou, and J. G. Schneider, "Docimological quality analysis of LLM-generated multiple choice questions in computer science and medicine," SN Computer Science, vol. 5, no. 5, Art no. 636, 2024.

M. Kaya, E. Sonmez, A. Halici, H. Yildirim, and A. Coskun, "Comparison of AI-generated and clinician-designed multiple-choice questions in emergency medicine exam: A psychometric analysis," BMC Medical Education, vol. 25, no. 1, Art no. 949, 2025.

A. Strugatski and G. Alexandron, "Applying IRT to distinguish between human and generative AI responses to multiple-choice assessments," in Proceedings of the 15th International Learning Analytics and Knowledge Conference, ACM, Mar. 2025, pp. 817-823.

M. Sabqat, R. A. Khan, M. Jawaid, and M. Sajjad, "Artificial intelligence meets item analysis (AI meets IA): A study of chatbot training and performance in detecting and correcting MCQ flaws," Pakistan Journal of Medical Sciences, vol. 41, no. 3, pp. 652-656, 2025.

Y. Shi, K. Yu, Y. Dong, and F. Chen, "Large language models in education: A systematic review of empirical applications, benefits, and challenges," Computers and Education: Artificial Intelligence, vol. 10, Art no. 100529, Jun. 2026.

S. Moore, H. A. Nguyen, T. Chen, and J. Stamper, "Assessing the quality of multiple-choice questions using GPT-4 and rule-based methods," in Responsive and Sustainable Educational Futures. Springer Nature Switzerland, 2023, vol. 14200, pp. 229-245.

S. Moore, E. Costello, H. A. Nguyen, and J. Stamper, "An automatic question usability evaluation toolkit," in Artificial Intelligence in Education. Springer Nature Switzerland, 2024, vol. 14830, pp. 31-46.

S. Guizani, T. Mazhar, T. Shahzad, W. Ahmad, A. Bibi, and H. Hamam, “A systematic literature review to implement large language model in higher education: issues and solutions,” Discover Education, vol. 4, no. 1, Art no. 35, 2025.

T. A. May, Y. K. Fan, G. E. Stone, K. L. K. Koskey, C. J. Sondergeld, T. D. Folger, J. N. Archer, K. Provinzano, and C. C. Johnson, "An effectiveness study of generative artificial intelligence tools used to develop multiple-choice test items," Education Sciences, vol. 15, no. 2, Art no. 144, 2025.

C. Isley, J. Gilbert, E. Kassos, M. Kocher, A. Nie, E. Brunskill, B. Domingue, J. Hofman, J. Legewie, T. Svoronos, et al., “Assessing the quality of AI-generated exams: A large-scale field study,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 40, no. 45, pp. 38626–38634, 2026.

R. B. Jabr and A. M. Azmi, "Knowledge-Aware Arabic question generation: A transformer-based framework," Mathematics, vol. 13, no. 18, Art no. 2975, 2025.

Y. Artsi, V. Sorin, E. Konen, B. S. Glicksberg, G. Nadkarni, and E. Klang, "Large language models for generating medical examinations: Systematic review," BMC Medical Education, vol. 24, no. 1, Art no. 354, 2024.

M. Kanık, "The use of ChatGPT in assessment," International Journal of Assessment Tools in Education, vol. 11, no. 3, pp. 608-621, 2024.

S. Al Faraby, A. Romadhony, and Adiwijaya, "Analysis of LLMs for educational question classification and generation," Computers and Education: Artificial Intelligence, vol. 7, Art no. 100298, Dec. 2024.

L. Goorts, R. Hollevoet, V. Xia, F. Cammaerts, and A. Güngör, "How do LLMs perform in the context of MCQs across different levels of thinking skills in a business education course at higher education? A comparison of ChatGPT, Gemini, and Copilot," Computers and Education: Artificial Intelligence, vol. 9, Art no. 100475, Dec. 2025.

J. Doughty, Z. Wan, A. Bompelli, J. Qayum, T. Wang, J. Zhang, Y. Zheng, A. Doyle, P. Sridhar, A. Agarwal, et al., "A comparative study of AI-generated (GPT-4) and human-crafted MCQs in programming education," in Proceedings of the 26th Australasian Computing Education Conference, ser. ACE 2024, ACM, Jan. 2024, pp. 114-123.

J. S. P. Baudin, "Assessing the psychometric properties of AI-generated multiple-choice exams in a psychology subject," Journal of Pedagogical Sociology and Psychology, vol. 7, no. 3, pp. 18–34, 2025.

M. Elzayyat, J. N. Mohammad, and S. Zaqout, "Assessing LLM-generated vs. expert-created clinical anatomy MCQs: A student perception-based comparative study in medical education," Medical Education Online, vol. 30, no. 1, Art no. 2554678, 2025.

Q. Zhao and M. Zhang, "Elimination-based reasoning with LLM for multiple-choice educational question answering," Journal of King Saud University Computer and Information Sciences, vol. 37, no. 7, Art no. 204, 2025.

H. Maeda, "Field-Testing multiple-choice questions with AI examinees: English grammar items," Educational and Psychological Measurement, vol. 85, no. 2, pp. 221-244, 2024.

H. Fawareh, "Evaluation of cloud computing for advancement LMS through different environments," International Journal of Advances in Soft Computing and its Applications, vol. 16, no. 3, pp. 125–148, 2024.

H. W. Awalurahman and I. Budi, "Automatic distractor generation in multiple-choice questions: A systematic literature review," PeerJ Computer Science, vol. 10, Art no. e2441, Nov. 2024.

L. Benedetto, S. Taslimipoor, and P. Buttery, "A survey on automated distractor evaluation in multiple-choice tasks," in Proceedings of the 20th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2025), Association for Computational Linguistics, Jul. 2025, pp. 55-69.

T. M. Haladyna and S. M. Downing, "A taxonomy of multiple-choice item-writing rules," Applied Measurement in Education, vol. 2, no. 1, pp. 37-50, 1989.

M. Tavakol and R. Dennick, "Making sense of Cronbach's alpha," International Journal of Medical Education, vol. 2, pp. 53-55, Jun. 2011.

J. Cohen, Statistical power analysis for the behavioral sciences. Routledge, May 2013.

J. R. Landis and G. G. Koch, "The measurement of observer agreement for categorical data," Biometrics, vol. 33, no. 1, pp. 159-174, 1977.

1857.image

Downloads

Additional Files

Key Dates

Received

30-03-2026

Revised

11-07-2026

Accepted

19-07-2026

Published

30-09-2026

Data Availability Statement

The main data used to reach the conclusions of this research are published in the paper and supplementary material. To further transparency and reproducibility, additional response matrices have been shared, which display the frequency distribution of student responses for each MCQ item.

Issue

Section

Original Article

How to Cite

[1]
O. D. Sarhan, “Hybrid Framework for Pre-Assessing Arabic Multiple-Choice Questions Using ChatGPT and Traditional Item Analysis”, Al-Mustansiriyah J. Sci., vol. 37, no. 3, pp. 1–13, Sep. 2026, doi: 10.23851/mjs.v37i3.1857.

Similar Articles

21-30 of 323

You may also start an advanced similarity search for this article.