Automated Formative Assessment With Large Language Models: Design, Validation, and Empirical Application in Higher Education
DOI:
https://doi.org/10.33423/jhetp.v25i6.8016Keywords:
higher education, large language models (LLM), formative assessment, automated grading, personalized feedback, AI in higher education, natural language processingAbstract
This study presents the design and validation of a formative assessment system that automatically grades and provides feedback on open-ended responses in higher education using large language models (LLMs). Implemented locally through the Ollama platform with a customized course corpus, it generates an “ideal response” based on instructor criteria and evaluates student answers accordingly. Semantic filters and sequential tolerance thresholds ensure alignment between AI and human grading.
A dataset of 450 exams (2022–2025) from three technical courses was used, with 360 for model optimization and 90 for validation. Performance metrics included mean difference, Pearson correlation, intraclass correlation coefficient (ICC), and Bland–Altman analysis. Results showed strong agreement (mean absolute difference = 4.5/100, r > 0.95, ICC ≈ 0.95) and negligible bias (±8/100).
The system generated detailed, pedagogically coherent feedback and effectively filtered irrelevant or ambiguous responses. Overall, it provided accurate and consistent assessment, greatly reducing grading time and enhancing formative learning while preserving instructor oversight.