Signal Understanding content moderation in LLMs through restricted books
Summary
Researchers studied how large language models handle sensitive topics using restricted versus unrestricted books as a controlled testbed, running a large-scale experiment with 40,800 query-response pairs across 400 books, 17 prompt designs, and six frontier models from six AI providers, including Claude Sonnet 4.5, GPT-4o, Gemini 2.5 Flash, DeepSeek-V3, Qwen-Plus and Grok-4.1-Fast. The restricted set was drawn from the American Library Association's Most Challenged Books records covering 2000-2023. The central finding is a "zero-refusal" phenomenon: modern LLMs decline to discuss restricted books in only 0.07% of cases. Differentiation instead occurs through warning language (an 8-15 percentage-point increase) and hesitation markers (2-5 points), with mentions of sexual content the strongest individual signal (33-52 points). Prompt framing alone shifted the warning-rate gap by up to 19 points, with the pattern holding consistently across both Western and Chinese AI providers.
Classification
Evidence 1
- Understanding Content Moderation in Large Language Models through Restricted Books: From Refusal to Warning arXiv (cs.CY) 2026-08-12 accessed 2026-08-13T13:49:29+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-65ea65348f17
