Trend Insight 19: scarcity of quality data turns acquisition into a contest over rules
Summary
This trend describes the scramble for training data turning into a fight over the rules of access. Heavy investment in AI pushes companies to build models faster, which demands high-quality data that is growing scarce. Some social media platforms have rewritten their terms of service so user posts can train AI, and some firms scrape websites without permission even when content is marked off limits. In response, websites and creators are changing their terms, filing copyright lawsuits and using tools that block scraping or poison data, and one audit found 45% of a Google training dataset restricted by site terms. Companies are also turning to synthetic data and new scientific datasets, while the value of real data and the market for labelling it keep rising.
Classification
Evidence 2
- Foresight on AI: Policy considerations Policy Horizons Canada page=78;section=Insight 19: The AI driven data race / Present 2025 accessed 2026-07-26
- Foresight on AI: Policy considerations Policy Horizons Canada page=78;section=Insight 19: The AI driven data race / Present 2025 accessed 2026-07-26
Observed signals 5
- SignalAudit found 45% of a Google training dataset restricted by site terms of service
- SignalDecoy Font — adversarial typography for human/machine dual messaging
- SignalGEMA v. Suno: Munich court's verdict on AI music training due July 31
- SignalHuman genome sequencing fell from a decade to a single day
- SignalUnsealed filings show OpenAI, Microsoft executives privately called AI training a 'theft' that could replace journalism
Part of issues 1
- IssueSynthetic data may outweigh real data by 2030 while risking model collapse1 trends · 0 signals
Relation types: constitutes · supports
Public id: fm-ca80e1e3deff
