Pew Research dropped a number this week that should make anyone who works with words sit up. Of every web page published since ChatGPT arrived in November 2022, more than one in three now carries signs that a machine wrote or heavily edited it. The finding lands just weeks after Cloudflare reported that bot traffic has overtaken human traffic, and the two numbers stack into an uncomfortable picture of what the internet is quietly turning into.

What Pew actually found
Pew pulled nearly half a million English-language pages from the Common Crawl archive, reaching all the way back to 2021, and ran them through Open Pangram, an open-weight AI detection model. In a random sample of 10,000 pages collected this July, about 10% showed significant signs of AI authorship. That slice sounds small, until you remember how old most of the web is.
Filter out everything published before ChatGPT, and the picture sharpens fast. Over a third of newer pages, 35% by Pew’s count, carry those same signals. The older pages in the sample could not have been written by a model, so they drag the overall average down. Pew is careful to frame this as directionally correct. Detection tools misclassify individual documents all the time, and the researchers say so. But run the numbers across hundreds of thousands of pages and the pattern separates cleanly enough to trust the trend.
And the trend is climbing. The share of .com pages showing AI authorship went from about 1.1% in 2022 to 9.4% by the start of 2026. This is not a weird corner of the web anymore. It is a growing slice of the entire thing.
Where the machine text lives
Location matters more than the average. Around 10% of .com pages showed AI authorship, against 4.6% of .org and roughly 1% each of .edu and .gov. That split is basically a map of where writing is a cost instead of a purpose. A commercial domain exists to be found. It gets crawled, ranked, and monetized, and search-optimized text is the cheapest thing a language model can produce by a wide margin.
Educational and government pages move slower, partly because the people who write them answer for accuracy and partly because nobody is mining them for ad revenue. The signal is a useful heuristic: the stronger the commercial incentive, the more likely a page was generated rather than written.
The tells, and the irony in them
Pew catalogued what gives machine text away, and at scale those tells have become more common across the web. Em dashes appear about twice as often as they did in 2023. Oxford commas are up 63%. Words that models reach for, like “delve,” “interplay,” and “testament,” have more than doubled. And the “it’s not just X, it’s Y” construction has nearly tripled.
There is a layer of irony here for anyone who edits. None of these traits is objectively wrong. Em dashes and Oxford commas are legitimate style choices, and I use an em dash now and then without thinking twice. The problem is that models lean on them so consistently that they turned into statistical fingerprints. The same habits that make AI prose easy to read are the habits that mark it as machine-made.
The recursion problem nobody asked for
The deeper issue is what happens when models start writing for other models to read. A growing share of new text gets generated, crawled, and fed right back into training data. Earlier research this year put AI articles at close to half of newly published pieces, even though human-written work still dominates search results and AI citations today. The old default assumption of the web, that a published page was probably written by a person, is quietly going away.
That recursion threatens everything downstream. If tomorrow’s models train on a web thick with machine text, they inherit its blandness and its errors. It is the same loop that keeps coming up whenever we talk about AI content, and it is why original, verifiable human work becomes more scarce and more valuable at once, even as new models of publishing treat words as raw material, the way Amazon has started using books to feed its AI.
What this means for readers
On a practical level, treat newer pages with a bit more skepticism. A page published this year was probably touched by a model somewhere. That does not make it wrong automatically, but it does mean vibes are no longer a reliable guide. Cross-check the claim, open the actual source, and stay alert for manufactured authority, since fake think tanks have started fooling even AI chatbots.
This is also a fresh reminder about trust. we use AI more while trusting it less, and that gap is widest exactly where reliable information matters most. Detection and labeling are improving, and text watermarks are starting to help with provenance, but none of that restores the default credibility the web used to hand out for free. A reader now has to do more of the work, and the tools that should be making verification easier are the same tools generating the noise that needs verifying.
What this means for people who write
For writers, editors, and publishers, the takeaway is sharp: originality is the differentiator now. If a model can produce most of the content in your niche at a fraction of the cost, your job is to own the part it cannot fake, real experience, real testing, real opinion, real accountability. That is what E-E-A-T has been circling for years, and the Pew numbers make it concrete.
I run this blog the way I do because that is the edge. Nobody comes to Bleuken for a generic recap they could grab from a chatbot. They come because a person actually ran the tool, tested the setup, or thought it through from a Filipino developer’s messy reality. That becomes more valuable, not less, as the machine-written noise keeps rising.
So where does that leave us? The web crossed a line where you can no longer assume another human wrote what you are reading. The fix is not to distrust everything, and it is not to trust everything either. It is to build the habit of checking, and to keep making the case for work that only a person can vouch for. That habit is cheap. The cost of losing it is the whole point of the web.