Signal Study Finds LLMs' Ability to Track "Where Information Came From" Breaks Down Over Conversation Memory
Summary
Researchers tested whether large language models possess "reality monitoring" — the capacity to distinguish their own generated output from information a user provided — across two experiments involving six different LLMs. Under minimal memory demands, models showed near-ceiling accuracy at identifying self-generated content, correctly attributing their own prior outputs. However, once an episodic delay was introduced into the conversation, that accuracy reversed into a fragile advantage for identifying externally provided items instead. Providing feedback exposed two distinct failure patterns: in some models, internal and external attribution judgments swapped entirely, while in others accuracy improved but confidence became decoupled from correctness — a dissociation invisible to existing benchmarks. The researchers found this pattern was linked to active parameter count rather than aggregate model size across the tested systems. They argue that as AI systems take on more autonomous, multi-turn roles, evaluating what a model knows is not sufficient — tracking where that knowledge came from may matter just as much, given links between reality-monitoring failure and hallucination or confabulation in humans.
Classification
Evidence 1
- arXiv (cs.AI, cs.CL, cs.CY, q-bio.NC) 2026-07-27 accessed 2026-07-28T15:00:27+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-7b92b4c39f81