Signal Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models
Summary
Researchers conducted a pre-registered audit examining whether a general-purpose helpfulness rubric can distinguish direct answer-giving from pedagogical guidance in LLM tutoring. Within each of three tutor model bases, they compared conversational and pedagogical policies instantiated with the same underlying model, paired with a fixed weak simulated student, while deterministic detectors measured answer leakage and independent student work in the next turn. Claude Opus 4.8 served as the frozen, condition-blind primary judge, and after its scores were fixed, GPT-5.6 Sol was used for a prospective post hoc robustness audit of 1,179 confirmatory answer-phase tutor turns. On the primary tutor base, the two policies did not differ significantly in helpfulness scores but were perfectly rank-separated under the pedagogy rubric. Across the two judges, pedagogy contrasts retained their direction where detected, but the helpfulness ordering was judge-dependent, reversing between judges on two of the three tutor bases. Separately, the study found that answer-revealing turns were consistently followed by less independent student work across every base, a result the authors describe as judge-invariant by construction. The authors conclude that general-purpose helpfulness is not a reliable pedagogy signal in this controlled setting and recommend pairing pedagogy-targeted rubrics with deterministic process measures.
Classification
Evidence 1
- arXiv (cs.CL/cs.AI/cs.CY) 2026-07-30 accessed 2026-08-01T03:50:22+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-ce5c02fc6e9e