Signal Beyond Score Accuracy: Examining the Diagnostic Quality of LLM-Generated Structured Assessment in Higher Education
Summary
A study using the JorGPT dataset of 3,041 student responses to 50 open-ended computer science questions, scored by both human instructors and three commercial large language models, identified systematic gaps between the apparent and actual diagnostic quality of AI-generated grading and feedback in higher education. The sub-dimension rubric scores produced by the LLMs were found to be highly correlated with each other, meaning they provide redundant rather than genuinely independent diagnostic information. The textual feedback generated by the models rarely identified students' underlying misconceptions, detecting them in only 5 to 7% of cases compared to 15 to 31% for human teachers, functioning more as a coverage checklist than a true diagnostic tool. The tone of AI-generated feedback also remained uniformly positive regardless of the actual quality of the student's response, lacking the severity modulation seen in human-written feedback. Grading accuracy varied significantly by knowledge domain, with procedural topics proving most reliable for the models. The authors argue the findings provide empirically grounded guidance on which aspects of LLM-based grading can be trusted and which still require continued human oversight.
Classification
Evidence 1
- Beyond Score Accuracy: Examining the Diagnostic Quality of LLM-Generated Structured Assessment in Higher Education arXiv 2026-10-07 accessed 2026-10-08T04:25:29+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-bc66a6859daf
