Does Google penalise AI content? No. It never has, it says so in writing, and the sentence doing the work has been sitting on a public Google URL since February 2023.
What Google penalises is content made at scale to game rankings, and it'll punish that whether a person typed it or a model did. The distinction sounds like a technicality and it's the whole thing. It means the question worth asking about your archive has nothing to do with the tool that produced it.
Below is what Google states, what the largest recent study found across 331,000 pages, the three policies that do bite, and a practical list of what to do about it. There's also a section on Anthropic's watermark, which started marking Claude's output this month and which almost nobody has read carefully enough to notice it doesn't apply to most of the models people are actually using yet.
What the evidence says
The first five come from an Ahrefs study of a million SERP positions published on 27 July 2026, sample of roughly 150,000 pages with enough text to analyse and 80,861 tracked in Search Console. The last is from Google's own spam policy documentation. Both checked on 11 August 2026.
What Google Actually Says
The primary source is a Search Central post from 8 February 2023, written by Danny Sullivan and Chris Nelson for the Search Quality team. Its subheading answers the whole question on its own: Rewarding high-quality content, however it is produced.
The body is blunter still. Google says its focus on the quality of content rather than how it got produced is a guide that's served the company for years, and then reaches for a comparison most people arguing about this have never read. About a decade before that post there were worries about a rise in mass-produced but human-written content. Google points out that nobody thought the sensible answer was to ban human writing.
Using AI does not give content any special gains. It is just content. If it is useful, helpful, original, and satisfies aspects of E-E-A-T, it might do well in Search. If it does not, it might not.
That's a close paraphrase of Google's own published answer to whether AI content ranks. It isn't a loophole and it isn't grudging tolerance. It's the position, published, never retracted, and consistent with everything the company has said since.
The Sentence Everybody Skips
Here's where most articles on this query stop, and where they mislead. The same post says that using automation, AI included, to generate content mainly to manipulate search rankings breaches the spam policies. Both halves are load-bearing.
Danny Sullivan has since said so himself, in unusually direct terms. His complaint was that the industry took the first half of the statement and ignored the second, and that it curdled into a general belief that Google doesn't care about AI. That was the wrong message, in his words, and what people should have taken from it was to ask whether the content is helpful.
Worth being accurate about who's speaking. Sullivan stopped being Google's Search Liaison on 1 August 2025 and is now a director inside Google Search, and nobody replaced him in the liaison role. The 2023 guidance still stands under his name.
The 331,000-Page Study
Policy statements tell you the rule. They don't tell you whether the rule survives contact with the ranking systems. For that, the best evidence available is an Ahrefs study by Ryan Law and Xibeijia Guan, published on 27 July 2026.
They pulled a million pages from the top ten positions across 100,000 searches, kept the roughly 300,000 their crawler already held, and ran AI detection on the 150,000 with enough text to make detection mean anything. Then they tracked 80,861 of those pages through Search Console across a year. That's a serious sample by the standards of this argument, where most published claims rest on a few dozen pages.
The Gradient Is Almost Flat
The headline finding is the one nobody expects. Average AI content of pages ranking first is 27.1%. Average AI content of pages ranking tenth is 30.9%.
Ahrefs, 27 July 2026. Bars are drawn to the real proportions, which is the point of the figure: the whole of page one fits inside four percentage points.
Under four percentage points separate the top of page one from the bottom of it. If Google ran an AI penalty of any strength that gradient would be a cliff. It's a slope you could push a pram up.
| Measure | Finding | Reading |
|---|---|---|
| AI share, position 1 vs 10 | 27.1% vs 30.9% | No meaningful penalty gradient |
| Top-ranking pages at 100% AI | 5.3% | Rare, but it happens |
| Top-three results under 50% AI | 82.2% | The winners are mostly blended |
| Indexation, low AI vs very high AI | 49.28% vs 40.35% | A real gap, well short of a ban |
| Impressions, low and moderate vs high | 2 to 3x | Where the damage actually shows |
All five from the same Ahrefs dataset, July 2026.
Where the Data Does Bite
Read the bottom two rows and a different story turns up, and any honest answer to this question has to include it. Pages with low or moderate AI content took two to three times the impressions of heavily-AI pages. Indexation rates were nine percentage points apart.
So heavily-AI pages do worse. They're just not doing worse because a system spotted AI and docked them for it. Law's own conclusion is that Google isn't trying to punish AI-generated content and is leaning on the same hallmarks of quality it always has.
That matches what the numbers look like from the inside. A page that's 95% unedited model output tends to be thin, tends to say the average thing, tends to contain nothing checkable and tends to have no position. Those four properties were ranking liabilities long before anybody could generate them at speed.
Where Governed Output Lands on the Same Scale
The Ahrefs study sorts pages into four buckets by detected AI percentage. Low under 20%, moderate 20 to 50%, high 50 to 80%, very high at 80% or above. Low and moderate are the buckets that took the impressions.
So here's where I sit on that scale, since it's what a client asking does Google penalise AI content actually wants to know. Every article on this site is drafted by a model. By production method they're 100% AI. Run them through the detectors clients use and they come back at 2% to 10%, which isn't just inside the low bucket, it's at the floor of it.
| Bucket | Detected AI | How it performed |
|---|---|---|
| Low | Under 20% | 2 to 3x the impressions, 49.28% indexation, where my pages score |
| Moderate | 20 to 50% | Also in the winning half. 82.2% of top-three sits under 50% |
| High | 50 to 80% | Falling away |
| Very high | 80% and above | 40.35% indexation, a third of the impressions |
Bucket definitions and performance from Ahrefs, July 2026. The score for my own pages is my measurement, described below.
Which Is Not the Contradiction It Looks Like
Two paragraphs up I said detection tools are unreliable and the wrong thing to worry about. I still say it. A detector can't tell you who wrote something, and I wouldn't defend a page on the strength of its score.
It matters commercially anyway, for two reasons that have nothing to do with whether the tool is any good. Clients run them, and a procurement conversation ends quickly at 96%. And the Ahrefs study measured its buckets with the same class of tool, so whatever the score really represents, it's the variable that correlates with two to three times the impressions.
What a Low Score Actually Measures
Not human authorship. Detectors read surface properties. Sentence-length uniformity, vocabulary predictability, rhythm. A low score means those properties are missing, and they're missing here because a script strips them out of every draft.
That sounds like gaming the metric until you look at what the same pass does to the page. You can't strip out uniform rhythm and average vocabulary without putting something in their place, and the only things available are specifics, positions and numbers. The score falls as a side effect of the work that makes the page worth reading.
Which is this whole article restated from the other end. People asking whether Google penalises AI content are worrying about the wrong half of the process, because the search engine isn't scoring the tool that produced the words, it's scoring what the finished page contains, and the edit that happens to satisfy a detector is the same edit that gives a reader a reason to stay.
The 3 Things Google Does Penalise
In March 2024, alongside a core update, Google added three spam policies. Not one of them mentions AI in its definition. All three catch the behaviour people mistake for an AI penalty.
| Policy | What it covers | Who it catches |
|---|---|---|
| Scaled content abuse | Many pages generated primarily to manipulate rankings rather than help users | Anyone publishing volume with no value added, by any method |
| Site reputation abuse | Third-party content hosted on a site mainly to borrow that site's ranking signals | Publishers renting out subfolders, the practice known as parasite SEO |
| Expired domain abuse | Buying a lapsed domain and repurposing it to inherit its ranking history | Domain flippers |
The scaled content policy is the one that matters here, and its defining clause deserves quoting exactly. Google's wording is that the practice is typically focused on creating large amounts of unoriginal content that provides little to no value to users, no matter how it's created.
Five words, doing an enormous amount of work. The policy is explicitly method-blind. A team that hires forty offshore writers to produce four hundred near-identical location pages sits inside the same policy as a team that generates them, and always did.
What the March 2024 Enforcement Actually Hit
Enforcement followed immediately and it wasn't subtle. Reporting at the time counted hundreds of sites deindexed within days, and a study by Originality.ai of 79,000 sites found manual actions on around 2% of them, covering more than twenty million monthly visits.
Those sites weren't deindexed for using a model. They were deindexed for publishing at volume with nothing behind it, which you can do with a model, a content mill, or a spreadsheet of spun synonyms.
The Statistic That Gets Misquoted
One number from that period appears in almost every article on this subject: 100% of deindexed sites showed signs of AI content. It gets presented as proof of an AI penalty. Two things are wrong with that.
First, the 100% figure came from a hand-checked sample of fourteen sites, not from the 79,000. A finding across fourteen cases is a reason to look harder, not a law. Second, Originality.ai sells an AI detector, so it has a direct commercial interest in the conclusion that AI content attracts penalties. That doesn't make the study wrong. It does mean reading it with the same care you'd bring to a vape company's research on vaping.
Every site publishing junk at scale in 2024 was using AI, because that was the cheapest way to publish junk at scale in 2024. The correlation is real and the causal story is backwards.
That distinction is the sort of thing I'd want an editor to catch in my own copy, and it's exactly why a fact-checking step has to be its own protocol rather than something bundled into a general read-through.
Why the Myth Persists
If the documentation is public and the data is unambiguous, it's worth asking why people searching does Google penalise AI content find so many pages telling them yes. Four reasons, and only one of them is honest confusion.
Somebody Is Selling the Answer
A big share of the writing on this question comes from companies selling AI detectors, agencies selling human copywriting, or tools selling AI-humanising. Each has a commercial reason for you to believe the tool you use decides your rankings. None of them is lying exactly. They're answering a question next door to yours.
And the tell is consistent. Where an article's answer is yes, scroll to the bottom and read what the publisher does for a living. It settles the disagreement faster than reading the argument does.
The Correlation Really Is There
This is the honest one. Sites did get hit in 2024, those sites were full of machine output, and anyone watching it happen drew the obvious conclusion. Correlation that strong feels like proof.
What it missed is that the cheapest way to publish worthless pages at volume changed in 2023, that everybody in the business of publishing worthless pages at volume switched to it more or less at once, and that for about eighteen months afterwards the spammers and the heavy AI users were close enough to the same set of people that a penalty aimed at either would have looked exactly like a penalty aimed at the other.
Nobody Reads the Second Half
Google's own guidance has the yes and the no in consecutive paragraphs, and the industry quoted whichever half suited it. Sullivan's public frustration about that is the closest thing to an official correction anyone has issued, and it got a fraction of the coverage the original post did.
The fourth reason is simply that does Google penalise AI content is a more interesting article than the true answer, which is that quality still decides it and always has. Only one of those gets shared.
Does It Answer the Question?
Strip out the policy language and one test survives. Somebody typed a query. Did your page answer it, and did it answer it better than the page above yours?
Everything else is downstream of that. Helpful content isn't a system you optimise for by hitting a word count and a heading structure. It's the plain question of whether a reader who arrived with a problem leaves with it solved, and if you can answer that honestly you can stop worrying about detection altogether.
The Test I Run Before Publishing
Read the piece as the person who searched, not as the person who commissioned it. If your page sends them back to the results to click the next one, it doesn't matter what produced it.
Three failures account for nearly everything I catch at that stage. The page answers a slightly different question to the one asked. The page answers the right question but takes eleven paragraphs to get there. Or it answers the question and gives the reader nothing they couldn't have guessed.
What to Do About It
None of this is theoretical. The gap between output that ranks and output that doesn't is a short list of production habits, and they're cheap.
Eight things that measurably improve the output
- Demand facts and figures. Ask for a number, a date, a named source or a mechanism in every section, then delete the sections that cannot produce one.
- Feed it your own knowledge first. Positioning, customer, the thing competitors get wrong. No model has access to it, and it is the only reliable source of non-average output.
- Appraise it with a second model. Paste the draft into a different LLM and ask what is weak, unsupported or generic. A model marking another model's homework is unsentimental about it.
- Verify every claim yourself. Second-model appraisal catches weak arguments and confidently invents corrections. Treat its output as a list of things to check.
- De-AI the prose. Ban the vocabulary tells, break uniform sentence rhythm, cut summary closers. Run it as a script so it happens whether or not anyone remembers.
- Optimise properly. Match the depth of whoever is ranking now, cover the questions the SERP shows, get the internal links right.
- Take a position. Balanced on every question is the clearest signal of unmanaged generation, and the easiest thing on this list to fix.
- Put a name on it. Somebody signs the piece, and that somebody read it.
The working list, in the order I run it. Steps one and two are the ones teams skip, and they are the two that decide whether the output is average.
Why the Second-Model Appraisal Works
It's the cheapest step on the list and the one people find hardest to believe. Ask the model that wrote a draft to critique it and you get politeness. Ask a different model, with nothing invested in the text, to name the three weakest claims and it'll name them.
The reason is boring rather than magical. A fresh model has no context telling it the draft is good, so it reads the text the way a reader would rather than the way an author does. A second opinion with no ego, available in seconds, and I've never once run it without changing something.
Two cautions. It'll invent problems as readily as it finds them, so every correction it proposes needs checking against a source. And it can't tell you whether the piece is worth publishing, which is still the half of the job that doesn't automate.
Claude Now Watermarks Its Output, and That Is Fine
There's a new fact in the argument as of this month, and it's worth understanding properly before somebody presents it to you as a threat.
Anthropic has started embedding an imperceptible watermark straight into text Claude generates. It applies to models launched on or after 2 August 2026, across the API, Claude, Claude Code and the rest, and it travels with the text when somebody copies and pastes it. Images and other supported files get a C2PA mark instead. The driver is the EU AI Act transparency code, and Anthropic applied it worldwide rather than just in Europe.
The cutoff is a launch date rather than a usage date, and that's easy to misread. Opus 5, the model behind this pipeline, launched on 24 July 2026, so it falls inside the transition period the AI Omnibus agreement gives systems already on the market, which runs to 2 December 2026. Nothing here is marked yet. Anthropic hasn't published a detector either, so nobody can test a page against a real reader at the moment. I've gone through all of that in detail, including what a real editing pass does to a mark once there is one.
What the Mark Proves and What It Does Not
Anthropic is careful about this and the distinction matters commercially. A mark says content may have been processed by Claude. It doesn't confirm where the content came from, it doesn't establish that Claude wrote it first, and the text may have changed since. Heavy editing, paraphrasing or translation can make the mark unreadable, and the absence of a mark proves nothing at all.
So it's a processing signal, not a verdict on authorship. That's a far weaker claim than the one people will read into it, and the weakness is deliberate.
Why This Is a Compliance Feature and Not a Ranking Risk
Think about what would have to be true for a watermark to hurt your rankings. Google would need to be reading a vendor-specific mark, treating its presence as a quality signal, and demoting pages on that basis. That would contradict its published position, its spam policy wording and the behaviour visible across 331,000 pages.
It'd also be a strange system to build, because the mark only says a tool touched the text. Worse, it'd run backwards. The mark weakens as a draft gets edited, so the pages most likely to still carry a clean one are the pages nobody worked on.
For a regulated publisher the watermark is closer to an asset. If you're an operator in a market where a regulator can ask how a page was produced, a provenance standard that answers mechanically beats an affidavit. Same argument as gating a page before it publishes rather than apologising afterwards.
Google Asks You to Disclose Anyway
The 2023 guidance has a line in it almost nobody quotes, and it settles a question I get asked constantly. Google says AI or automation disclosures are useful for content where someone might wonder how was this created, and suggests adding them wherever that would be reasonably expected.
That's an invitation rather than a requirement, and it's why every article here carries a note at the foot explaining how it was made. The same guidance says accurate author bylines matter where a reader might wonder who wrote something, and that giving an AI a byline is probably not the way to follow the recommendation.
Both instincts point the same way. A human name at the top, an honest note at the bottom, and nothing pretending to be something it isn't.
What This Means If You Publish at Scale
Volume is the variable everybody worries about and on its own it's the wrong one. The scaled content policy has no page-count threshold in it. What it has is a purpose test, and that test asks whether the pages exist to help somebody or to occupy a search result.
A thousand pages that each answer a genuinely distinct question aren't scaled content abuse. A hundred pages that are the same page with a place name swapped are, and were long before anybody had a model to swap them with.
The Practical Threshold
Here's the version I give clients. If you can't say what a page offers that the others in the set don't, that page is a liability, and the number of those is your actual exposure.
Run that test across an archive and the answer comes back in clusters rather than evenly. One template, generated once, applied three hundred times. That's the shape of the risk, and it's a production decision rather than a tooling one.
Questions I Get Asked
Does Google Penalise AI Content in 2026?
No. The guidance hasn't changed since February 2023, the spam policies added in 2024 are explicitly method-blind, and the largest study to date found under four percentage points of difference in AI content between first and tenth position. There's no penalty because there's no rule to penalise against.
Can Google Detect AI Content?
Partly, and it's the wrong thing to worry about. Detection tools of every kind throw false positives on careful human writing and false negatives on edited machine writing. More to the point, Google has no stated interest in classifying production method, so a detection capability would be a system with nothing to do.
Does the Spelling Change the Answer?
No. Search for penalise or penalize and you get the same policy, and this page is in British English because the rest of the site is.
Will the Claude Watermark Get My Site Flagged?
There's no mechanism by which it could. The mark is readable by Anthropic's own detection, it indicates processing rather than authorship, and no search engine has said it reads vendor watermarks as a ranking input. Treat it as provenance plumbing.
Should I Say a Page Was Written With AI?
Where a reader would reasonably wonder, yes, and Google says so directly. On a personal essay or a first-hand review the question is live. On a currency converter it isn't.
Is Human Content Still Better?
Governed content is better. The comparison that matters isn't human against machine, it's governed against ungoverned, and plenty of human writing fails every test above.
How Much AI Is Too Much?
The data gives you a rough answer rather than a rule. Pages under 50% detected AI take 82.2% of top-three placements, and heavily-AI pages see two to three times fewer impressions. Those numbers describe how much editing went in. They aren't a quota to hit.
What About AI Overviews and LLM Citations?
Same answer, different surface. The available research finds most URLs cited in AI Overviews also rank in the organic top ten, so the work that earns one tends to earn the other. Chopping your content into bite-sized fragments because you reckon an LLM prefers it is a specific thing Google has told people not to do.
How This Piece Was Made
Every figure in this piece traces to either Google's own documentation or the Ahrefs study, both read directly and both linked. The Originality.ai finding is included with its sample size and its conflict of interest stated, because that number is quoted everywhere without either. The Anthropic watermark detail comes from Anthropic's own support documentation, published this month.
One figure here is mine. The 2% to 10% detector range is my own measurement of my own pages, and you should read it as testimony rather than as a study. It carries the obvious interest: I sell the process that produces it. Run one of these articles through whichever detector your procurement team trusts and you will have a number that does not depend on taking my word for it.
An earlier version of this note stated flatly that this page carries a Claude watermark. It almost certainly doesn't, and not for the reason I gave at the time: the model behind it launched nine days before the marking rule applied. Claiming it as fact was the error the rest of the article warns about, and a reader's question caught it rather than any script. What the page definitely carries is my name, which is the part that matters. Tell me if you think any of it is wrong.