Over a third of web pages published since ChatGPT's debut show clear signs of AI generation or heavy machine editing, according to a recent analysis by Pew Research. Evaluating nearly half a million English-language pages from the Common Crawl archive across five years using Open Pangram classifiers, researchers found that while an unfiltered July 2024 crawl hovered around 10% synthetic text due to historical backlog, isolating post-launch material pushed machine authorship to 35%.

Commercial domains are spearheading this influx. Pew recorded AI authorship rates on `.com` domains at roughly ten times those of `.edu` and `.gov` sites (both around 1%), while `.org` settled at 4.6%. Beyond domain metrics, the study logged unmistakable spikes in algorithmic stylistic fingerprints, such as structural em-dash crutches and Oxford comma density.

For AI engineering leads and enterprise labs, this synthetic contamination directly threatens pretraining pipelines. As human-authored ground truth dries up, ingesting uncurated web dumps into next-generation foundation models risks systemic model collapse—accelerating reasoning degradation, compounding hallucinations, and exponentially driving up data scrubbing budgets. Combined with Cloudflare's finding that automated traffic has surpassed human users, we are rapidly approaching an absurd equilibrium: automated scrapers ingesting synthetic slop generated by other bots to train models that will yield progressively worse outputs.

Generative AILarge Language ModelsMachine LearningAI Safety