Future Monitor 한국어

Signal Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models

Summary

Researchers conducted a pre-registered audit examining whether a general-purpose helpfulness rubric can distinguish direct answer-giving from pedagogical guidance in LLM tutoring. Within each of three tutor model bases, they compared conversational and pedagogical policies instantiated with the same underlying model, paired with a fixed weak simulated student, while deterministic detectors measured answer leakage and independent student work in the next turn. Claude Opus 4.8 served as the frozen, condition-blind primary judge, and after its scores were fixed, GPT-5.6 Sol was used for a prospective post hoc robustness audit of 1,179 confirmatory answer-phase tutor turns. On the primary tutor base, the two policies did not differ significantly in helpfulness scores but were perfectly rank-separated under the pedagogy rubric. Across the two judges, pedagogy contrasts retained their direction where detected, but the helpfulness ordering was judge-dependent, reversing between judges on two of the three tutor bases. Separately, the study found that answer-revealing turns were consistently followed by less independent student work across every base, a result the authors describe as judge-invariant by construction. The authors conclude that general-purpose helpfulness is not a reliable pedagogy signal in this controlled setting and recommend pairing pedagogy-targeted rubrics with deterministic process measures.

Classification

Main topicAI & Computing
Secondary topicsEducation & Generations
Region menusGlobal
Impactscope:global
Time horizon4-10 years (2026-08-01)
Last updated2026-08-01T03:55:37.555820+00:00

Evidence 1

Part of trends 0

No objects.

Directly linked issues 0

No objects.

Public id: fm-ce5c02fc6e9e