A public dashboard observing signals, trends and issues.
SubscribeLogin한국어
Latest observation
2026-10-08
Public objects
4434
Build time
2026-10-08 19:44 KST
The Futures

Signal Study compares instruction- vs. example-driven AI content moderation

Summary

A study assesses whether foundation models can apply content moderation policy consistently in practice. The study compares two approaches. One gives the model explicit written instructions describing the policy, and the other trains it on labeled examples of the policy being applied. The authors argue that as moderation rules grow more complex, enforcing them consistently becomes an increasingly serious challenge. With that concern in mind, the researchers test whether foundation models actually have the capacity to turn written policy into consistent moderation decisions. The findings bear on the methodological reliability of using large language models for platform governance and content-policy enforcement at scale.

Classification

Region menusGlobal
Impactscope:global
Time horizon0-3 years (2026-09-11)
Last updated2026-09-25 22:32 KST

Evidence 1

Part of trends 0

No objects.

Directly linked issues 0

No objects.

Public id: fm-8490f237701f