Signal EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners
Summary
Unggi Lee, Sookbun Lee, Yeil Jeong, Eunjoo Lee, Minchul Shin, and Hoilym Kwon published a paper on arXiv (cs.CY) on August 4, 2026, introducing EduClaw-Bench, a benchmark addressing the gap in evaluating AI tutor agents over sustained relationships rather than single sessions. The benchmark places an agent tutor in a continuous 30-day relationship with a simulated learner grounded in knowledge tracing (KT), whose concept-mastery levels, driven by a KT model trained on real-student data, determine its answers and are probed for learning gain across 55 scenarios. Each agent is scored on three primary axes (learning gain, responsiveness, and helpfulness) and two curriculum-design axes (Gagné and Rosenshine instructional models), with helpfulness and curriculum axes judged by a cross-family panel of three LLM judges. Evaluating 10 agent adapters across three base-model tiers yielded two findings unavailable to single-tier, single-session evaluation: tutoring quality belongs jointly to the base model and the agent harness rather than either alone, and almost no combination sustains good tutoring across the full 30-day horizon. A calibration check (ECE=0.049) and a live-classroom field study confirmed that the simulated learner and its measurements track real classroom behavior.
Classification
Evidence 1
- arXiv (cs.CY) 2026-08-04 accessed 2026-08-05T02:34:16+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-5228d339b6c6