Explainable Machine Learning Berbasis Simulasi Respons Untuk Analisis Kualitas Soal

Authors

  • Hidayattullah Universitas Amikom Yogyakarta Author
  • Abdul Madjid Hasanuddin Universitas Amikom Yogyakarta Author
  • Murniati Universitas Amikom Yogyakarta Author
  • Frankgling Nusa Universitas Amikom Yogyakarta Author

DOI:

https://doi.org/10.61805/fahma.v24i3.240

Keywords:

explainable machine learning, XGBoost, SHAP, analisis butir soal, efektivitas pengecoh

Abstract

Multiple-choice item quality analysis is essential for educational assessment quality assurance, yet conventional psychometric analysis relies on examinee response data that are often unavailable during item development. This study develops an explainable machine learning framework based on simulated responses to analyze multiple-choice item quality using 10,962 CommonsenseQA items. The dataset serves as a computational testbed due to its consistent five-option format, available answer keys, large scale, and plausible distractors, rather than as a substitute for validated educational item banks. Responses from 200 virtual respondents per item were probabilistically simulated based on answer-key weighting and surface-level distractor plausibility. Four psychometric indicators—difficulty index, discrimination index, distractor effectiveness, and distractor entropy—combined with five surface linguistic features were used to train XGBoost to classify item quality as poor, moderate, or good. Five-fold cross-validation yielded a weighted F1-score of 0.9987 ± 0.0011 and accuracy of 0.9987 ± 0.0011. Feature importance and SHAP identified distractor entropy and discrimination index as dominant predictors. The near-perfect performance is not interpreted as external predictive validity because the target and primary predictors derive from the same simulation mechanism. This framework provides a transparent, auditable proof-of-concept for preliminary item-bank evaluation before empirical testing.

Downloads

Download data is not yet available.

References

T. M. Haladyna, S. M. Downing, and M. C. Rodriguez, “A Review of Multiple-Choice Item-Writing Guidelines for Classroom Assessment,” Applied Measurement in Education, vol. 15, no. 3, pp. 309–333, Jul. 2002, doi: 10.1207/S15324818AME1503_5.

F. B. Baker and S.-H. Kim, The Basics of Item Response Theory Using R. Cham: Springer International Publishing, 2017. doi: 10.1007/978-3-319-54205-8.

G. Kurdi, J. Leo, B. Parsia, U. Sattler, and S. Al-Emari, “A Systematic Review of Automatic Question Generation for Educational Purposes,” Int. J. Artif. Intell. Educ., vol. 30, no. 1, pp. 121–124, Mar. 2020, Accessed: Sep. 30, 2026. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S1560429226006967

B. Settles, G. T. LaFlair, and M. Hagiwara, “Machine Learning–Driven Language Assessment,” Trans. Assoc. Comput. Linguist., vol. 8, pp. 247–263, Dec. 2020, doi: 10.1162/tacl_a_00310.

L. A. Ha, V. Yaneva, P. Baldwin, and J. Mee, “Predicting the Difficulty of Multiple Choice Questions in a High-stakes Medical Exam,” in Proceedings of the Fourteenth Workshop on Innovative Use of NLP for Building Educational Applications, Stroudsburg, PA, USA: Association for Computational Linguistics, 2019, pp. 11–20. doi: 10.18653/v1/W19-4402.

S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Proceedings of the 31st International Conference on Neural Information Processing Systems, in NIPS’17. Red Hook, NY, USA: Curran Associates Inc., 2017, pp. 4768–4777.

S. M. Lundberg et al., “From local explanations to global understanding with explainable AI for trees,” Nat. Mach. Intell., vol. 2, no. 1, pp. 56–67, Jan. 2020, doi: 10.1038/s42256-019-0138-9.

L. Benedetto, G. Aradelli, P. Cremonesi, A. Cappelli, A. Giussani, and R. Turrin, “On the application of Transformers for estimating the difficulty of Multiple-Choice Questions from text,” in Proceedings of the 16th Workshop on Innovative Use of NLP for Building Educational Applications, J. Burstein, A. Horbach, E. Kochmar, R. Laarmann-Quante, C. Leacock, N. Madnani, I. Pilán, H. Yannakoudakis, and T. Zesch, Eds., Online: Association for Computational Linguistics, Apr. 2021, pp. 147–157. [Online]. Available: https://aclanthology.org/2021.bea-1.16/

M. Byrd and S. Srivastava, “Predicting Difficulty and Discrimination of Natural Language Questions,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), Stroudsburg, PA, USA: Association for Computational Linguistics, 2022, pp. 119–130. doi: 10.18653/v1/2022.acl-short.15.

S. AlKhuzaey, F. Grasso, T. R. Payne, and V. Tamma, “Text-based Question Difficulty Prediction: A Systematic Review of Automatic Approaches,” Int. J. Artif. Intell. Educ., vol. 34, no. 3, pp. 862–914, Sep. 2024, doi: 10.1007/s40593-023-00362-1.

V. Yaneva et al., “Findings from the First Shared Task on Automated Prediction of Difficulty and Response Time for Multiple-Choice Questions,” in Proceedings of the 19th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2024), E. Kochmar, M. Bexte, J. Burstein, A. Horbach, R. Laarmann-Quante, A. Tack, V. Yaneva, and Z. Yuan, Eds., Mexico City, Mexico: Association for Computational Linguistics, Jun. 2024, pp. 470–482. [Online]. Available: https://aclanthology.org/2024.bea-1.39/

A. Talmor, J. Herzig, N. Lourie, and J. Berant, “CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Stroudsburg, PA, USA: Association for Computational Linguistics, 2019, pp. 4149–4158. doi: 10.18653/v1/N19-1421.

M. Yasunaga, H. Ren, A. Bosselut, P. Liang, and J. Leskovec, “QA-GNN: Reasoning with Language Models and Knowledge Graphs for Question Answering,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Stroudsburg, PA, USA: Association for Computational Linguistics, 2021, pp. 535–546. doi: 10.18653/v1/2021.naacl-main.45.

C. Hao, M. Xie, and P. Zhang, “ACENet: Attention Guided Commonsense Reasoning on Hybrid Knowledge Graph,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Stroudsburg, PA, USA: Association for Computational Linguistics, 2022, pp. 8461–8471. doi: 10.18653/v1/2022.emnlp-main.579.

M. Harwell, C. A. Stone, T.-C. Hsu, and L. Kirisci, “Monte Carlo Studies in Item Response Theory,” Appl. Psychol. Meas., vol. 20, no. 2, pp. 101–125, Jun. 1996, doi: 10.1177/014662169602000201.

L. Benedetto, G. Aradelli, A. Donvito, A. Lucchetti, A. Cappelli, and P. Buttery, “Using LLMs to simulate students’ responses to exam questions,” in Findings of the Association for Computational Linguistics: EMNLP 2024, Stroudsburg, PA, USA: Association for Computational Linguistics, 2024, pp. 11351–11368. doi: 10.18653/v1/2024.findings-emnlp.663.

T. Chen and C. Guestrin, “XGBoost,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, New York, NY, USA: ACM, Aug. 2016, pp. 785–794. doi: 10.1145/2939672.2939785.

T. Puthiaparampil and M. Rahman, “How important is distractor efficiency for grading Best Answer Questions?,” BMC Med. Educ., vol. 21, no. 1, p. 29, Dec. 2021, doi: 10.1186/s12909-020-02463-0.

A. A. Rezigalla et al., “Item analysis: the impact of distractor efficiency on the difficulty index and discrimination power of multiple-choice items,” BMC Med. Educ., vol. 24, no. 1, p. 445, Apr. 2024, doi: 10.1186/s12909-024-05433-y.

P. Razavi and S. Powers, “Estimating item difficulty using large language models and tree-based machine learning algorithms,” Int. J. Artif. Intell. Educ., vol. 36, no. 3, p. 100015, Sep. 2026, doi: 10.1016/j.ijaied.2026.100015.

Downloads

Published

30-09-2026

How to Cite

Explainable Machine Learning Berbasis Simulasi Respons Untuk Analisis Kualitas Soal. (2026). FAHMA : Jurnal Informatika Komputer, Bisnis Dan Manajemen, 24(3), 349-358. https://doi.org/10.61805/fahma.v24i3.240

Similar Articles

21-30 of 63

You may also start an advanced similarity search for this article.

Most read articles by the same author(s)

1 2 3 4 5 6 7 8 9 10 > >>