Signal Cultural Awareness is Represented but Not Decoded: Tracing Mythological Knowledge across 18 Open-Source LLMs
Summary
Iaroslav Chelombitko, Ekaterina Chelombitko, and Mika Hämäläinen published a paper on arXiv (cs.CY) on August 3, 2026, examining why open-source large language models reliably recognize deities such as Zeus, Jupiter, and Thor but are far less consistent in recognizing counterparts from Finnish, Slavic, Egyptian, or Chinese mythology. Using a parallel cross-cultural set of Thompson-motif mythological entities, the researchers applied linear probing, logit lens, activation patching, and output extraction across 18 open-source LLMs from 8 different architecture families. They found that the residual stream inside the models clearly distinguishes between cultures, performing well above a simple name-string baseline, meaning cultural information is represented internally. However, the decoder stage collapses culturally specific tokens onto dominant-tradition equivalents, so the failure occurs at the readout stage rather than in representation. The researchers also found that asking the same question in a culture's native language versus in English produces failure patterns that cluster within each language but diverge across languages, suggesting the decoder is gated by the prompt language. The team released a per-entity evaluation framework and citation-anchored ground truth dataset covering all 18 models.
Classification
Evidence 1
- arXiv (cs.CY) 2026-08-03 accessed 2026-08-05T02:02:05+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-08b3251ece51