e-ISSN: 2618-6586 Open Access · Peer-Reviewed JETOL on DergiPark
JETOL — Journal of Educational Technology & Online Learning
Publisher
Gürhan Durak
Publication Model
Periodical (January – May – September)
Status
Open for Submissions

Optimizing AI-based assessment in history education: The impact of prompt engineering on scoring performance in a morphologically rich language

Research Article

Download PDF DOI: 10.31681/jetol.1882892
Published in
Volume 9, Issue 3 (2026)
Pages
374–396
Publication Date
September 30, 2026
Submission Date
February 5, 2026
Acceptance Date
August 19, 2026
Subjects
Educational Technology and Computing

Abstract

This study investigates the effectiveness and prompt sensitivity of leading large language model (LLM)-based AI tools (ChatGPT, Gemini, Claude, Deepseek) in grading open-ended exam questions in an undergraduate history course (Atatürk’s Principles and History of Revolution), compared to human evaluators. A comprehensive dataset comprising 72 distinct open-ended responses (collected from 24 students) was scored by both the course instructor and seven AI models using five different prompts with varying levels of detail. These prompts ranged from a basic zero-shot instruction to progressively more structured designs: expected topic headings, a weighted criterion-based rubric, partial coverage with language proficiency, and relative (norm-referenced) scoring. Correlation and statistical analyses (paired samples t-test, Wilcoxon signed-rank test) revealed that while models showed low agreement with the human evaluator when using unstructured, basic prompts (zero-shot), they achieved high agreement (r > .80) when provided with structured prompts and explicit rubrics, particularly in the cases of Gemini and Claude. The findings highlight that for effective AI-based assessment, prompt design and the definition of criteria are more critical than model selection in mitigating "generosity bias." Generosity bias here denotes the models’ systematic tendency to assign higher scores than the human rater; the qualitative evidence indicates that it stems primarily from the models rewarding fluent, lengthy, and well-structured answers even when their factual content is incomplete, a tendency that explicit rubrics substantially reduced. Designed as an exploratory case study, these results provide significant empirical evidence regarding the optimization of AI as an assistive assessment tool, specifically within the context of Turkish, a morphologically rich language.

Keywords

  • AI in Education
  • Natural Language Processing
  • Morphologically Rich Languages
  • History Education
References (37)
  1. American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. Washington, DC: American Educational Research Association.
  2. Attali, Y., & Burstein, J. (2006). Automated essay scoring with e-rater V.2. The Journal of Technology, Learning and Assessment, 4(3).
  3. Badger, E., & Thomas, B. (1992). Open-ended questions in reading. Practical Assessment, Research & Evaluation, 3(4), 1991-1993.
  4. Bektaş, M., & Kudubeş, A. A. (2014). Bir ölçme ve değerlendirme aracı olarak yazılı sınavlar. Dokuz Eylül Üniversitesi Hemşirelik Fakültesi Elektronik Dergisi, 7(4), 330-336.
  5. Brucks, M., & Toubia, O. (2025). Prompt architecture induces methodological artifacts in large language models. PloS one, 20(4), e0319159. https://doi.org/10.1371/journal.pone.0319159
  6. Brown, H. D. (2004). Language assessment, principles and classroom practices. USA: Longman.
  7. Burrows, S., Gurevych, I., & Stein, B. (2015). The eras and trends of automatic short answer grading. International Journal of Artificial Intelligence in Education, 25(1), 60-117. https://doi.org/10.1007/s40593-014-0026-8
  8. Capdehourat, G., Amigo, I., Lorenzo, B., & Trigo, J. (2025). On the effectiveness of LLMs for automatic grading of open-ended questions in Spanish. arXiv preprint arXiv:2503.18072. https://doi.org/10.48550/arXiv.2503.18072
  9. Carlson, N., & Burbano, V. (2025). The use of LLMs to annotate data in management research: Foundational guidelines and warnings. Strategic Management Journal. Advance online publication. https://doi.org/10.1002/smj.70023
  10. Cooney, T. J., Sanchez, W. B., Leatham, K., & Mewborn, D. S. (2004). Open-ended assessment in math: A searchable collection of 450+ questions. “Open-ended Assessment in Math.”
  11. Dikli, S. (2006). An overview of automated scoring of essays. The Journal of Technology, Learning and Assessment, 5(1).
  12. Farray Rodríguez, J., Fernández-García, A. J., & Verdú, E. (2026). Grading open-ended questions using LLMs and RAG. Expert Systems, 43(1), e70174. https://doi.org/10.1111/exsy.70174
  13. Flodén, J. (2025). Grading exams using large language models: A comparison between human and AI grading of exams in higher education using ChatGPT. British educational research journal, 51(1), 201-224. https://doi.org/10.1002/berj.4069
  14. He, J., Rungta, M., Koleczek, D., Sekhon, A., Wang, F. X., & Hasan, S. (2024). Does Prompt Formatting Have Any Impact on LLM Performance?. arXiv preprint arXiv:2411.10541. https://doi.org/10.48550/arXiv.2411.10541
  15. Holmes, W., Iniesto, F., Anastopoulou, S., & Boticario, J. G. (2023). Stakeholder perspectives on the ethics of AI in distance-based higher education. International Review of Research in Open and Distributed Learning, 24(2), 96-117. https://doi.org/10.19173/irrodl.v24i2.6089
  16. Jauhiainen, J. S., & Guerra, A. G. (2024). Evaluating Students’ Open-ended Written Responses with LLMs: Using the RAG Framework for GPT-3.5, GPT-4, Claude-3, and Mistral-Large. arXiv preprint arXiv:2405.05444. https://doi.org/10.48550/arXiv.2405.05444
  17. Jonsson, A., & Svingby, G. (2007). The use of scoring rubrics: Reliability, validity and educational consequences. Educational research review, 2(2), 130-144. https://doi.org/10.1016/j.edurev.2007.05.002
  18. Kasneci, E., Sessler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., ... & Kasneci, G. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and individual differences, 103, 102274. https://doi.org/10.1016/j.lindif.2023.102274
  19. Kubiszyn, T., & Borich, G. D. (2024). Educational testing and measurement. John Wiley & Sons.
  20. Latif, E., & Zhai, X. (2024). Fine-tuning ChatGPT for automatic scoring. Computers and Education: Artificial Intelligence, 6, 100210. https://doi.org/10.1016/j.caeai.2024.100210
  21. Linn, R., & Miller, M. (2005). Measurement and Assessment in Teaching (9th ed.). Merrill-Prentice Hall.
  22. Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741-749. https://doi.org/10.1037/0003-066X.50.9.741
  23. Milakis, E. D., Argyrakou, C. C., Loukisas, T., Vangeli, D. G., Katsarou, M. C., & Lalou, A. (2024). Comparative analysis of large language models as AI assistants for educational Assessment: A case study in steam education. International Journal of Computing and Artificial Intelligence, 5(2), 51-55. https://doi.org/10.33545/27076571.2024.v5.i2a.97
  24. Morjaria, L., Burns, L., Bracken, K., Levinson, A. J., Ngo, Q. N., Lee, M., & Sibbald, M. (2024). Examining the efficacy of ChatGPT in marking short-answer assessments in an undergraduate medical program. International Medical Education, 3(1), 32-43. https://doi.org/10.3390/ime3010004
  25. Moskal, B. M., & Leydens, J. A. (2000). Scoring rubric development: Validity and reliability. Practical assessment, research, and evaluation, 7(1).
  26. OECD. (2023). OECD digital education outlook 2023: Towards an effective digital education ecosystem. OECD Publishing.
  27. Page, E. B. (1966). The imminence of... grading essays by computer. The Phi Delta Kappan, 47(5), 238-243.
  28. Pinto, G., Cardoso-Pereira, I., Monteiro, D., Lucena, D., Souza, A., & Gama, K. (2023, September). Large language models for education: Grading open-ended questions using chatgpt. In Proceedings of the XXXVII brazilian symposium on software engineering (pp. 293-302). https://doi.org/10.1145/3613372.3614197
  29. Popham, W. J. (2020). Classroom Assessment: What Teachers Need to Know (Ninth Edition). University of California, Los Angeles.
  30. Seo, H., Hwang, T., Jung, J., Kang, H., Namgoong, H., Lee, Y., & Jung, S. (2025). Large Language Models as Evaluators in Education: Verification of Feedback Consistency and Accuracy. Applied Sciences (2076-3417), 15(2). https://doi.org/10.3390/app15020671
  31. Shermis, M. D., & Burstein, J. (Eds.). (2013). Handbook of automated essay evaluation: Current applications and new directions. Routledge.
  32. The jamovi project (2024). jamovi. (Version 2.6) [Computer Software]. Retrieved from https://www.jamovi.org.
  33. UNESCO. (2023). Guidance for generative AI in education and research. UNESCO.
  34. Williamson, D. M., Mislevy, R. J., & Bejar, I. I. (Eds.). (2006). Automated scoring of complex tasks in computer-based testing. Mahwah: Lawrence Erlbaum Associates. https://doi.org/10.4324/9780415963572
  35. Yan, L., Sha, L., Zhao, L., Li, Y., Martinez-Maldonado, R., Chen, G., Li, X., Jin, Y., & Gašević, D. (2024). Practical and ethical challenges of large language models in education: A systematic scoping review. British Journal of Educational Technology, 55(1), 90-112. https://doi.org/10.1111/bjet.13370
  36. Yavuz, F., Çelik, Ö., & Yavaş Çelik, G. (2025). Utilizing large language models for EFL essay grading: An examination of reliability and validity in rubric-based assessments. British Journal of Educational Technology, 56(1), 150-166. https://doi.org/10.1111/bjet.13494
  37. Zawacki-Richter, O., Marín, V. I., Bond, M., & Gouverneur, F. (2019). Systematic review of research on artificial intelligence applications in higher education–where are the educators?. International journal of educational technology in higher education, 16(1), 39. https://doi.org/10.1186/s41239-019-0171-0

This article is published under the CC BY 4.0 license.