/ SEO  ·  August 21, 2026  ·  7 min read

Pew counted how much of the web is written by AI. The number changes what your content has to do.

Pew Research Center analysed roughly 490,000 pages from Common Crawl and published the result on August 20, 2026: about 10 percent of pages now show significant signs of AI authorship, up from around 1 percent in early 2021, and more than a third of pages published since ChatGPT's release. The tells (em dashes at twice the 2023 rate, Oxford commas up 63 percent, 'delve' and 'testament' more than doubled) are also leaking into human writing, which makes detection useless as evidence and turns them into a credibility tax instead. What that means for a small business publishing anything.

By Rushil Shah
SEOSmall BusinessAI

Everyone has an opinion about how much of the web is now machine-written. Yesterday Pew Research Center published an actual count.

Their data labs team sampled roughly 490,000 English-language webpages from Common Crawl archives spanning January 2021 to July 2026 and ran them through an open AI-detection model, scoring a page as showing meaningful signs of AI authorship above a set threshold. The headline result, published August 20, 2026:

  • About 10 percent of pages in the July 2026 crawl show significant signs of AI authorship.
  • Among pages carrying a publication date after ChatGPT’s release, the figure is more than a third.

The trend line is the part worth sitting with. On .com domains the share was roughly 1 percent in early 2021, 2.76 percent by mid-2023, 4.7 percent in July 2024, 6.64 percent in July 2025, and about 10 percent in July 2026. That is not a spike. That is a steady annual doubling-ish of a share that is now large enough to change what the average page on the internet is.

It also splits sharply by domain type. Roughly 10 percent for .com, 4.6 percent for .org, and around 1 percent for .edu and .gov. Commercial sites moved first and moved hardest, which will surprise nobody who has watched a competitor’s blog go from four posts a year to four posts a week.

The tells, and why you should not trust them

Pew also measured the linguistic fingerprints that have become associated with generated text. Compared with 2023:

  • Em dashes appear about twice as often.
  • Oxford commas are up 63 percent.
  • Specific vocabulary, including “delve,” “testament,” and “interplay,” more than doubled.
  • Negative parallelism, the “it’s not just X, it’s Y” construction, nearly tripled.

Two things follow, and most commentary gets only the first one.

The first is the obvious one: these are now recognisable, and readers recognise them. The em dash in particular has gone from a piece of punctuation to a social signal, which is genuinely unfair to everyone who has been using it correctly since school.

The second is the one that matters more. Pew notes that as more generated content appears online, these tells have become more common across webpages generally. Human writers read the web, absorb its cadence, and reproduce it. So the markers are contaminating the control group. Every year they become weaker evidence of what produced a page and stronger evidence only of when it was written.

Pew is explicit about the limits of the method: “AI detection models aren’t perfect, they sometimes misclassify individual documents that were written by humans as including signs of AI authorship, and vice versa.” The finding is a population-level one. It is reliable across half a million pages and worthless applied to any single page.

Which means: do not run an AI detector on a freelancer’s work and treat the output as proof. It is not proof. It is a probability estimate from a model that has been trained on a moving target, applied to a sample size of one. People are losing contracts and getting accused over these scores, and the study everyone is citing as evidence explicitly says the tool does not support that use.

Why this matters commercially, and it is not the reason you think

The intuitive worry is “will Google penalise my AI-written content.” That is the wrong frame, and it has been for a while. Google’s position, restated in its guidance for generative AI features last updated July 10, 2026, is not about how a page was produced. It is about whether the page is commodity content: material that recycles what everyone else has already said. Its instruction is direct, do not just repeat what others on the internet have already said.

The economics are what changed, and they changed underneath the SEO question.

Producing a competent, correct, unremarkable 1,200-word article about your industry used to cost a few hundred dollars and half a day. It now costs approximately nothing and takes about ninety seconds. Anything whose production cost falls to zero has a market value that falls to zero with it. The 35 percent figure for recent pages is what that looks like from the outside: an enormous and growing volume of pages that are individually fine and collectively worthless, because there is no reason to read any particular one of them.

That has three practical consequences for a small business.

Volume has stopped working as a strategy. If your plan for organic visibility is to publish more, you are entering a race where the marginal cost of your competitors’ next article is nothing. You will not win it. Both the March 2026 core update and the May one that followed rewarded specificity and first-hand experience and punished thin aggregation, which is the algorithmic expression of the same economics.

Reading like a machine now costs you trust, whoever wrote it. This is the underrated one. Your prospective client has spent the last year reading generated text, has learned the shape of it, and will pattern-match your services page against that shape in about four seconds. If your page opens with a paragraph restating its own heading, lists three benefits in a tricolon, and closes by summarising what it just said, the reader’s conclusion is that nobody was home, and it does not matter that you wrote every word yourself on a Sunday. You are paying a credibility tax on someone else’s output.

The scarce thing is now the specific thing. A language model can produce an article about kitchen renovation costs. It cannot produce the sentence “the two-week lead time on cabinet hardware is the thing that actually blows up a Leaside reno schedule, not the countertops.” That sentence requires having done the work. Everything a model cannot know about your business is, by definition, what is now worth publishing: your real prices, your actual timelines, the failure mode you see over and over, the job you turned down and why, the neighbourhood-specific detail, the photograph of the thing you built.

What to do with your own writing

Not a style guide. A short editing pass that takes ten minutes and does most of the work.

  1. Find the first concrete claim. A number, a date, a price, a place name, a named tool. If the first one appears in paragraph five, the first four paragraphs are throat-clearing and can go.
  2. Cut the paragraph that restates the heading. Generated structure loves an introductory sentence that tells you what the section is about. The heading already did that.
  3. Ask what only you could have written. If every sentence in a section could appear on a competitor’s site with the business name swapped, delete the section or replace it with the specific version.
  4. Put a real number in. Prices, timelines, dimensions, failure rates, how many you have done. A page with numbers reads as first-hand because it usually is, and it is also the format both search engines and assistants find easiest to lift and cite.
  5. Read it aloud. Every construction that makes you wince is a construction a reader was going to wince at silently.
  6. Then check the punctuation, last. If you naturally write with em dashes, this is an unfair moment to be you, and you may reasonably decide to swap them for commas or a full stop while the association is this strong. Style, not ethics.

Where we stand on this

Worth being direct, since we are a studio that both uses these tools and writes about them.

We use AI assistance in our work, including on client projects, and we have written about where it earns its keep and where it does not. The standard we hold ourselves to on this blog is not “no model touched this.” It is that every factual claim traces to a primary source we actually opened, that the specifics come from work we have actually done, and that a named human is accountable for the whole thing. That is a standard about verification and accountability, which are the parts that were always doing the real work. “Written by a human” was never the guarantee anyone thought it was; there was an enormous amount of worthless human-written content on the internet in 2019.

The practical upshot is the same either way. As the share of pages that are competent, correct, and interchangeable climbs past a third of everything new, the value of being a primary source goes up, not down. That is true for classic ranking, and it is more true for the AI answers we wrote about in the piece on Google’s new AI search controls, where the system has to choose a handful of sources to cite and quote rather than ten to list.

If your site is currently full of content that could have been written about anyone, that is a fixable problem and usually a fast one, because the raw material is already in your head. Tell us what you actually know and we will help you turn it into pages worth citing.

● Taking new projects

Have something that needs shipping?

One call. Thirty minutes. You leave with an honest read on scope, timeline, and price, whether we're the right fit or not.