Somebody is spending two hours right now making a good draft worse. Shortening the long sentences. Salting in a few contractions. Breaking a paragraph in half. Then they paste it into a detector, get 82% AI, and go round again, and this time the paragraph with the actual figure in it comes out, because numbers apparently look machine-written.
That's the loop. And the thing sitting underneath it is that the tool they're fighting isn't finding anything. It's guessing, based on whether the words are the boring ones a model would have picked. The company with more training data than anyone couldn't make that work and pulled its own version off the market.
Then a fortnight ago the question changed shape, because for the first time there's an actual mark. Claude models launched from 2 August write an invisible watermark into their text, worldwide, no opt-out. That one isn't guessing, and it doesn't care how varied your sentence lengths are.
What an AI Detector Actually Does
A detector reads your finished text and guesses at the odds a model wrote it. There's no mark in there to find. It's doing statistics on word choice: how predictable each word is given the ones before it, and how much that predictability moves around.
Human writing lurches. A six-word sentence, then a forty-word one, then a word you didn't see coming. Model output runs smoother, because a model keeps picking the likely next word, so the detector hunts for smoothness and calls it machine. Which is why every guide gives you the same list, and why the whole thing falls over the moment you point these tools at properly crafted AI writing.
The Numbers the Detector Companies Don't Lead With
OpenAI built its own AI detector in 2023. It caught just 26% of AI-written text and gave false positives for 9% of human-written content, so OpenAI did the right thing and pulled the tool. The people behind the algorithm couldn't detect their own work, and publicly admitted it.
The independent work is worse. A Stanford study ran seven widely used detectors over 91 TOEFL essays written by non-native English speakers and logged an average false-positive rate of 61.3%. Of those 91 essays, 89 got flagged as machine-written by at least one detector. Point the same tools at essays by American eighth-graders and they behave themselves.
The mechanism is the damning part. Those detectors weren't finding AI. They were finding a smaller vocabulary and simpler sentences, which is what writing in your second language looks like, and what a model looks like when nobody has trained it. Two unrelated things, one fingerprint.
When the researchers rewrote those essays with richer vocabulary, the false-positive rate fell by about four fifths. So yes, the tool can be beaten. That's not reassuring. A test you can talk out of its answer with better word choice was never measuring what it claimed to.
A detector guesses. A watermark knows. Every technique on every "beat the detector" list is aimed at the guess.
The Advice on Every Page-One Guide, and What It Costs
I read the four pages ranking for this query. They agree almost word for word: vary your sentence lengths, cut the mechanical transitions, drop the corporate vocabulary, use contractions in most sentences, add a personal aside, run it through a humanising tool at the end.
Most of It Is Just Good Editing
Here's the awkward part. Nearly all of that is correct, and I do most of it myself, for reasons that have nothing to do with detection. Uniform sentence length is bad writing. "Furthermore" at the head of a paragraph is a tic. A page with no first-person judgement in it is a page nobody needed.
The difference is the reason. Do it because uniform rhythm is dull and you get better prose. Do it to move a score and you're optimising against a number that calls six in ten human beings a machine.
Where It Turns Destructive
The specifics go first. A figure with a source attached is exactly the kind of precise, information-dense sentence that reads as machine-written to these tools, so stripping those out leaves you with a page that passes the detector and loses to any page that kept its numbers. It's the same mechanism as the policy that quietly deletes true things, wearing a different hat.
Then the cure grows its own tell. At full strength the de-AI checklist gives you clipped, article-stripped, telegram prose that no person writes either, which is why removing the tells adds new ones. You don't arrive at writing that sounds human. You arrive at a second machine sound.
And a humaniser is itself a language model. You're paying a second model to rewrite the first model's output so a third model scores it differently. Nothing in that sentence even makes logical sense when you really think about it. It's a terrible way to work.
Nobody Is Going to Demote You for This
Worth saying plainly, because it's the fear sitting underneath the whole question. There is no step inside a search engine that detects AI and docks you for it. Google's spam policy goes after content made at scale to game rankings and says nothing at all about whether a machine helped, so there's no test to fail here and no score you have to stay under.
Which is why the threshold people keep asking me for doesn't exist. The largest study on this read a million ranked pages and put average detected AI content at 27.1% for position one, against 30.9% for position ten. Under four points separating the top of page one from the bottom of it. A page sitting at 30% isn't flirting with a penalty. It's indistinguishable from the pages already winning.
What Changed on 2 August
The EU AI Act's Article 50 became enforceable on 2 August 2026. It tells providers of generative systems to mark synthetic output in a machine-readable format, and its code of practice asks for two techniques rather than one, something like signed tamper-evident metadata plus an imperceptible watermark. Get it wrong and the fine runs to €15 million or 3% of worldwide turnover.
Anthropic's answer landed on 11 August. Every Claude model launched on or after 2 August writes an invisible statistical watermark into its text, and Anthropic applied it globally rather than just to European traffic. Files get signed provenance metadata under the C2PA standard, the same one Adobe and Google use.
How the Mark Is Made
The method is a version of SynthID-Text, published by Google DeepMind in Nature in 2024. As a model writes it keeps choosing between words that would do about the same job. Grey or overcast. Begin or start. Normally a random number settles that, and the watermark swaps the randomness for a function of a secret key and the words just before it, so the choices still look unpredictable to you while carrying a pattern the key can read.
That's a different class of thing from everything above. A detector infers. A watermark is a signal somebody put there on purpose. You can argue with the first one. The second is either in the words or it isn't.
| Classifier detector | Statistical watermark | |
|---|---|---|
| What it reads | Your finished text, cold | A pattern placed during generation |
| What it measures | How predictable the wording is | Agreement with a key |
| Fooled by rewriting | Yes, cheaply | Only by replacing the words |
| False positives on human text | Documented, and high | Not applicable, nothing to find |
| Who can run it | Anyone, for free | Whoever holds the key |
One limit worth knowing. Confidence climbs with length, so the mark reads badly on short passages, where there weren't enough word choices to carry a pattern. A paragraph is a weak sample. An article is a strong one.
Does the Mark Survive a Real Editing Process?
That's a question about my own pages. I write with Claude Code against three documents: a standard the draft is written against, a fix protocol the rewrite runs against, and a read-only audit that scores the published page. Every piece goes through all three, then through me. So I put it to Google's Gemini in plain terms and asked whether the watermark reaches the published piece.
"No, your final published piece will not carry Claude's watermark. A statistical watermark relies on an unbroken chain of specific token probabilities generated directly by the model. The moment you introduce multi-stage editing, human revisions, and manual SEO adjustments, that mathematical chain snaps."
The Mark Lives in Words That Get Rewritten
The watermark rides on the model's choices between near-equivalent words, which is the exact material an editing pass spends its time replacing. Anthropic is direct about what follows: whether the mark can still be read depends on how long the text is and how heavily it was edited, and where human revision dominates the finished words, there's very little watermark material left.
A structural rewrite, a fact pass that changes what the sentences claim, a voice pass against a fifty-nine word blacklist and a specificity pass that puts a number in every paragraph don't leave many original choices standing. That isn't proofreading. That's most of the text.
The Model That Wrote It Predates the Rule
The second reason is a date. The watermark applies to Claude models launched on or after 2 August 2026. Opus 5, which Claude Code runs and which drafted this piece, launched on 24 July 2026. Nine days before the line. Under the AI Omnibus agreement reached in May, systems already on the market before 2 August have until 2 December 2026 to meet the marking requirement.
So a piece published here carries no watermark, and that second reason has about fifteen weeks left to run. The first one doesn't expire. Anthropic says it'll offer a detection API and hasn't shipped one, so nobody outside the company can test a live page yet, me included. When it ships I'll run it over this site and publish whatever comes back.
Three Documents, and Why They Make the Question Boring
What that system optimises for isn't undetectability. It optimises for whether a claim survives an interview: every figure checked in the session it was written, every comparative tested against the named alternative, every link pinged before it ships. The voice rules exist because uniform prose is dull, and they're written down so the same correction never has to be made twice. If a detector likes the result, that's a side effect.
| Detector advice | Why a standard already says it | Same? |
|---|---|---|
| Vary your sentence lengths | Uniform rhythm is dull to read | Yes |
| Cut "furthermore", "moreover" | The heading carries the transition | Yes |
| Drop corporate verbs | They describe no action | Yes |
| Use contractions everywhere | Voice is set by page type, not by quota | No |
| Add typos and false starts | Never. It is a lie about how the page was made | No |
| Remove precise figures | The opposite. A page with no number loses | No |
| Run a humanising tool | Never. A second model is not an editor | No |
Four of the seven diverge, and those four are the ones that would actively damage the page. That's what makes the framing dangerous rather than merely useless. It agrees with good practice just often enough to get believed on the rest.
What to Do Instead
Disclose. It sounds like the soft answer and it dissolves the problem, because a detector score only threatens you if the reader thinks you claimed otherwise. Every article here carries a note saying how it was made, and two clients have told me that note is why they got in touch.
Then do the work the detector is a bad proxy for. Check the claims in the session you write them. Put in the number, the date, the named body, the mechanism. Take a side, because a model with no opinion is the actual tell and no amount of sentence-length variance hides it. Google's spam policy goes after content made at scale to game rankings and says nothing about whether a machine helped, which I've pulled apart with the study figures in a separate piece on the penalty that doesn't exist. No detector sits between your draft and the index.
If a Detector Has Already Accused You
Plenty of people searching this phrase aren't trying to get away with anything. They wrote something themselves, a tool called it machine-written, and now they want a technique to make it go away. Don't rewrite the work to satisfy the tool. Rewriting destroys the only thing that answers it. Keep the version history, the notes, the search results you had open, the draft with the bad paragraph still in it. A document that shows its own construction over several hours beats any score.
Then make whoever is holding that score say what it means. A tool that flags six in ten essays by competent non-native speakers isn't evidence on its own, and the firm with the best view of how these models write pulled its detector rather than keep defending it. Ask what the false positive rate is on writing like yours. Very few institutions have an answer.
What None of This Solves
The watermark belongs to one company, and one company isn't the field. Text from other providers carries no equivalent mark today, so a negative result proves little. A positive one proves less than people will assume. It says a model was involved somewhere, which for most working writers in 2026 is a fact about their toolchain rather than a confession.
And if you're passing off unedited generation as work you did, none of this helps you. The detector isn't your problem. The problem is that the page has nothing in it only you knew, which any reader paying attention can see without a tool at all.
Beating the detector was always the wrong ambition. Produce better content and the detector stops being the thing you work around. Building the system that does that, reliably, at volume, is what I do.