Signal LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure
Summary
Modern language models are trained on heterogeneous web-scale text, which makes it hard to pin down or rule out prior exposure to related content within the training data. As a result, studying how knowledge and skills are acquired in a controlled way has been difficult. To address this, the authors built a pretraining corpus called LITTLECURRICULUM that matches the level of US elementary school material and explicitly leaves out concepts, facts and vocabulary beyond that level. The corpus totals 88 billion tokens. The researchers trained a model called LittleLearner on this pedagogically controlled corpus, allowing knowledge and skill acquisition to be studied where the scope of exposure is precisely known and bounded. The work is presented as controlled-science infrastructure for interpretability and learning-dynamics research rather than a deployable product.
Classification
Evidence 1
- LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure arXiv (cs.AI) 2026-08-13 accessed 2026-08-16T10:59:38+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-382fec306fba
