A public dashboard observing signals, trends and issues.
SubscribeLogin한국어
Latest observation
2026-10-08
Public objects
4434
Build time
2026-10-08 19:44 KST
The Futures

Signal Beyond Score Accuracy: Examining the Diagnostic Quality of LLM-Generated Structured Assessment in Higher Education

Summary

A study using the JorGPT dataset of 3,041 student responses to 50 open-ended computer science questions, scored by both human instructors and three commercial large language models, identified systematic gaps between the apparent and actual diagnostic quality of AI-generated grading and feedback in higher education. The sub-dimension rubric scores produced by the LLMs were found to be highly correlated with each other, meaning they provide redundant rather than genuinely independent diagnostic information. The textual feedback generated by the models rarely identified students' underlying misconceptions, detecting them in only 5 to 7% of cases compared to 15 to 31% for human teachers, functioning more as a coverage checklist than a true diagnostic tool. The tone of AI-generated feedback also remained uniformly positive regardless of the actual quality of the student's response, lacking the severity modulation seen in human-written feedback. Grading accuracy varied significantly by knowledge domain, with procedural topics proving most reliable for the models. The authors argue the findings provide empirically grounded guidance on which aspects of LLM-based grading can be trusted and which still require continued human oversight.

Classification

Secondary topicsAI & Computing
Region menusGlobal
Impactscope:global
Time horizon0-3 years (2026-10-08)
Last updated2026-10-08 14:45 KST

Evidence 1

Part of trends 0

No objects.

Directly linked issues 0

No objects.

Public id: fm-bc66a6859daf