Are AI detectors reliable? The record says no, and Claude's watermark does not fix it
2026-08-14
Anthropic now watermarks every word Claude writes. But AI detection false positives are what actually cost people their work, their grades and their reputations, and a watermark proves a machine was involved, not that it wrote anything.
Built on verified case files. The argument below leans on evidence The Internet Ninja validated against the public record and published in full, method included.
- FTC final order: the '98% accurate' AI Content Detector was an unmodified student model, and its own published data showed 53.2% on real-world text
- FTC v. Evolv: the AI scanner marketed as detecting all weapons missed a knife used in a school stabbing, and a federal court entered a permanent injunction
- FTC × IntelliVision: a 'zero bias, millions of faces' AI claim, measured against NIST
- Federal court bans Rite Aid from AI facial-recognition surveillance for five years after FTC alleges thousands of false-positive matches
- A $2.275M class settlement takes the algorithm's score away: Louis v. SafeRent and the five-year rollback of tenant-screening scores for voucher applicants
- A federal court fined a public defender $1,500 for a fake citation, then declined to say AI wrote it
- Kohls v. Ellison: a court throws out a Stanford misinformation expert's declaration after GPT-4o invented its citations
- Zalando's 2018 'algorithms replace 250 marketing jobs': the headline, and the hiring plan it left out
- Klarna's AI customer-service assistant: the 2024 numbers and the 2025 walk-back
- Moderna's company-wide ChatGPT Enterprise rollout: 750 custom GPTs, from a vendor case study
- Notion and Decagon: an AI support agent, 34% faster resolution and 2x deflection
Kurzgesagt has more than twenty million subscribers and hand-animates every frame. In late July 2026, YouTube’s automated systems decided one of their science videos was AI slop. The channel watched it become one of its worst performers despite healthy click-through and watch time, then went looking for the cause. Their own account: “YouTube’s automatic AI detection tools wrongly think that our very much human-made videos are AI Slop,” and it “started to choke our channel” source. YouTube acknowledged the error.
Four days later, Anthropic announced it would watermark every word Claude writes.
Almost everyone covered the second event. The first one is the story.
Three accusations, none of them about anyone using AI
Photographer Peter Yan posted a picture of Mount Fuji to Instagram and found it labelled “Made with AI”. He had removed a bin. “I did not use generative AI, only Photoshop to clean up some spots. This ‘Made with AI’ was auto-labeled by Instagram when I posted it, I did not select this option” source.
PetaPixel then ran the experiment that matters. Remove a tiny speck with Generative Fill and the label appears. Make the identical edit with the Spot Healing Brush, Content-Aware Fill or Clone Stamp and it does not, “despite these tools having exactly the same effect on the overall image.” Same photograph. Same change. A different menu item, and a different verdict about whether reality is real.
Then YouTube began quietly applying machine-learning sharpening to creators’ uploads without telling them. Guitarist Rhett Shull compared his source files against what viewers saw, and objected on a specific ground: “it looks AI-generated. I think that deeply misrepresents me and what I do and my voice on the internet” source. The platform made his real footage look synthetic. The reputational cost landed on him.
Three cases. In none of them did anyone pass off machine writing as their own. In all three a human made something, a machine ruled on its origin, and the machine was wrong. That is the grievance of the last three years, and a watermark does not touch it.
What the Claude watermark actually proves
Claude models launched on or after 2 August 2026 embed a statistical watermark in generated text, with C2PA signed metadata on generated image files. It applies “wherever Claude is offered, worldwide,” not only in the EU, though the EU AI Act’s Article 50 is the stated driver source source. It covers the API, Claude Code, and Claude through AWS, Google Cloud and Microsoft Foundry.
Anthropic’s own documentation draws the line more carefully than most of the coverage did. A detected mark “indicates that the content may have been processed by Claude.” It “provides a signal that content was processed by Claude, but is not fully conclusive.” The page says plainly that “Claude may not be the original author. People often use Claude to proofread, translate, summarize, or convert files.” And it lists what breaks detection: content that is “heavily edited, paraphrased, translated, or mixed into other writing,” and short passages source.
Read that against the three cases above. A student who used Claude to fix grammar carries the same mark as one who generated the whole essay. A writer who ran a paragraph through for tone carries the same mark as a content farm. The mark answers one question: was a machine involved at any point. That question stopped separating anything the moment AI was built into every editor, phone keyboard and search box.
Anthropic also declines to carry the compliance burden it created. Its guidance tells those building on Claude to “independently assess what Article 50 requires of your products and services.” The vendor marks the output. Whether the mark satisfies anyone’s legal obligation stays the deployer’s problem.
The reader was a choice, and the industry wrote down why two years ago
Most coverage treated the missing detector as a timeline problem. It is better read as the third of three answers the major labs have given to the same question, and the other two are on the record.
Google DeepMind published its text watermarking scheme, SynthID-Text, in Nature on 23 October 2024 source, and open-sourced the implementation: the watermarking code plus three detector variants, mean, weighted mean and Bayesian, under Apache 2.0, with a production version in Hugging Face Transformers source.
Read that repository closely and the limit of the openness is visible. Detection needs the watermarking configuration, and the Bayesian detector must be trained for each unique watermarking key source. Publishing the algorithm does not create a public read path. A statistical text watermark stays legible only to whoever holds the key, whichever lab ships it.
Google’s own products show where that lands. As of 17 August 2026 the SynthID Detector portal verifies “an image, video or audio file”, the check inside Gemini covers “an image, video or audio clip”, and the portal is still a waitlist for journalists and media professionals source. Text is in neither surface, more than a year after text support was announced as coming at the portal’s launch source.
OpenAI went further and published its reasons for not shipping at all. In an update dated 4 August 2024 it confirmed that “our teams have developed a text watermarking method that we continue to consider as we research alternatives”, then listed the objections. The method is “less robust against globalized tampering; like using translation systems, rewording with another generative model, or asking the model to insert a special character in between every word and then deleting that character”. And separately: “our research suggests the text watermarking method has the potential to disproportionately impact some groups. For example, it could stigmatize use of AI as a useful writing tool for non-native English speakers” source.
Both objections are the substance of this article, published by a competitor two years before Claude’s mark shipped. The circumvention list is nearly identical to the one in Anthropic’s own documentation. The harm to non-native speakers is what Liang and colleagues measured in Patterns.
OpenAI drew one more distinction worth keeping. It favoured cryptographically signed metadata because “unlike watermarking, metadata is cryptographically signed, which means that there are no false positives”, and because “while text watermarking has a low false positive rate, applying it to large volumes of text would lead to a large number of total false positives” source. That is the Vanderbilt arithmetic, restated as engineering by a lab that decided not to ship.
So the three answers are: publish the method and keep the read path private, decline to ship and say why, or ship the mark and defer the reader. Only the third puts a mark into public circulation that nobody in public can read.
This section describes surfaces shipped as of its capture date, not intentions. Any of these labs could release a text read path tomorrow, and the honest version of this claim is dated for that reason.
AI detection has already failed, and it is on the public record
Commercially
The Federal Trade Commission ordered Workado, LLC to stop marketing its AI Content Detector as 98 percent accurate. Independent testing put it at 53 percent on general-purpose content, because the underlying model had been trained and validated on academic writing only source. We hold that final order as a verified case file.
It is not an isolated file. The same shape, an AI system sold on a detection-accuracy claim that did not survive contact with a regulator, runs through Evolv’s AI weapons screening, IntelliVision’s zero-bias facial recognition claims, and Rite Aid’s facial recognition, which drew a five-year ban after misidentifying shoppers. In Louis v. SafeRent, algorithmic tenant screening produced a $2.275m class settlement. Four regulators and one federal court, all reaching the same finding: these systems do not perform as advertised, and the cost lands on the person being scored.
Mathematically
Researchers at the University of Maryland showed recursive paraphrasing drops watermark detection from a 99.8 percent true-positive rate to 9.7 percent with minimal quality loss, and demonstrated the reverse attack, making human text look machine-generated source. Their theoretical result is the durable one. Detector accuracy is bounded by the statistical distance between human and AI text, and that distance shrinks every time a model improves. Detection carries an expiry date set by the technology it chases.
Worse than fragile, it is forgeable. ICML work on watermark stealing found that by querying a watermarked model’s API to reverse-engineer the scheme, “for under $50 an attacker can both spoof and scrub state-of-the-art schemes previously considered safe, with average success rate of over 80%” source. Spoofing means stamping a legitimate model’s watermark onto text it never produced. A separate ICML paper proves strong watermarking impossible under reasonable assumptions, even when detection is kept private source. Evidence that can be forged for the price of lunch is not evidence. It is a way to frame people.
Psychologically
A 2026 study in the Journal of Science Communication tested what a disclosure label does to belief, and found a truth-falsity crossover. The label reduced the perceived credibility of correct information while increasing the perceived credibility of misinformation source. The proposed mechanism is a machine heuristic: a false claim reads as more objective once flagged as machine-made, while true but nuanced writing is penalised for reading cold. In that data the label is not merely uninformative about truth. It points the wrong way. The study surveyed 433 Chinese social-media users on science-communication texts, so the direction of the effect is the finding, not a universal constant.
On real people
Vanderbilt University did the arithmetic on its own students and disabled Turnitin’s AI detector in 2023. At a claimed 1 percent false-positive rate against the 75,000 papers submitted the prior year, “around 750 student papers could have been incorrectly labeled.” Its other stated reasons are worth repeating. Turnitin “gives no detailed information as to how it determines if a piece of writing is AI-generated,” and detectors “have been found to be more likely to label text written by non-native English speakers as AI-written” source. Within weeks of launch Turnitin conceded a higher-than-expected false-positive rate, putting the sentence-level figure near 4 percent source.
The bias is measured, not anecdotal. In Patterns, Liang and colleagues ran seven detectors over TOEFL essays written by non-native English speakers and over US eighth-grade essays. More than half the non-native essays were misclassified as AI-generated. Accuracy on native writing was near-perfect. They confirmed the mechanism causally: enrich the vocabulary of a non-native essay and the flags disappear source.
It costs people work. Kimberly Gasuras, a news reporter of twenty-four years, was removed from a freelance platform after a detector flagged writing she says she wrote. A second Ohio copywriter lost the client representing roughly 90 percent of his income over a 95 percent AI score, despite timestamped drafting history source.
In February 2026 a New York court reached the obvious conclusion. In Matter of Newby v. Adelphi University, a judge annulled a university’s finding that a student’s paper was 100 percent AI-generated, describing it as without valid basis, and ordered the disciplinary record expunged source. We are building that decision into a full case file against the slip opinion. Until it clears our own gate, treat this paragraph as reported, not verified.
What the label is actually for
Look at what a disclosure label does in the one place a statute spells it out. Singapore’s Online Safety (Relief and Accountability) Act 2025 creates an offence of inauthentic material abuse, covering AI-manipulated content that falsely depicts someone. The section’s own illustrations state that material carrying a prominent AI-generated label does not qualify as inauthentic material source.
The label is a safe harbour. Disclose, and you sit outside the offence. It was never built to tell a reader whether the underlying claim is true.
The EU arrives at the same place from the other direction. Article 50(4) requires deployers to disclose AI-generated text published to inform the public on matters of public interest, unless the content “has undergone a process of human review or editorial control and where a natural or legal person holds editorial responsibility” source. The escape hatch is not the absence of AI. It is a human taking responsibility. That is the whole answer, sitting in an exemption clause.
Every regime says label. None can say what a label is.
Article 50 became applicable on 2 August 2026. No harmonised European standard sits behind its marking requirement. The Commission operationalised it through a voluntary Code of Practice and non-binding guidelines instead.
India notified the strictest labelling rules outside China in February 2026, then dropped the specification. The October 2025 draft required a label covering at least 10 percent of visual surface area, or the first 10 percent of an audio file’s duration. The final rules replaced every number with “clearly and prominently” source.
China runs the only enforcement machinery operating at scale, and it is genuinely large. The Cyberspace Administration’s April 2026 campaign named inadequate implementation of AI-content labelling as an explicit target. Its first phase disposed of more than 14,000 AI products, cleared over six million items and actioned 26,000 accounts source. It works by takedown and delisting, in aggregate, not by fining a named company for a labelling failure.
Singapore ran the closest thing to a controlled experiment. It legislated specifically against election deepfakes, then held a general election. In the five days before Nomination Day, 73 AI-generated election videos appeared on TikTok. Eleven contained digitally manipulated visuals of actual candidates, the Prime Minister in seven of them, and only four were clearly labelled as AI. The government’s assessment was that none breached the law, and no corrective directions were issued over them source. A purpose-built statute, a real flood of synthetic candidate content, and nothing to prosecute, because the law targets realistic deception rather than AI use, and most of it did not clear that bar.
Spain’s enforcement law is still in committee. South Korea’s regime commenced in January 2026 under a grace period during which fines are deferred. Put the sweep together and one fact holds across every jurisdiction we examined. No regulator anywhere has fined anyone specifically for failing to label AI-generated content. The obligation is live in Europe. The machinery is real in China. The adjudicated penalty does not exist.
The systems that decide what gets seen ignore all of it
Google’s spam policy defines the risk as scaled content abuse: many pages generated “for the primary purpose of manipulating search rankings and not helping users.” Its listed example is “using generative AI tools or other similar tools to generate many pages without adding value for users” source. The trigger is scale and absence of value. It has never been whether AI touched the page.
Google’s provenance work sits somewhere revealing: under Search appearance, not ranking. The May 2026 announcement brings SynthID and C2PA verification to Search and Chrome as something a curious person can ask about one item at a time, through Lens, AI Mode and Circle to Search source. That is opt-in verification, not a filter that runs before something is surfaced.
Here we owe you a distinction most coverage collapses. Google has never said provenance signals affect ranking. Google has also never said they do not, unlike structured data, where it has repeatedly gone on record. Anyone writing “Google confirmed provenance does not affect ranking” is extrapolating. So are we, if we claim the opposite.
What is measured points the other way. Ahrefs examined a million AI Overview result pages and half a million cited URLs, and found cited pages ran 3.6 percent pure-AI and 87.8 percent mixed human and AI, against a web where roughly 72 percent of new pages already contain AI-assisted content source. Answer engines cite AI-assisted content at or above its base rate. They are not filtering it out.
One caution, because this number is going to travel. An SEO study published the day of Anthropic’s announcement reports that watermarked content ranks about five positions lower and earns half the AI citations. Its own methodology note concedes it “did not control for content quality beyond our own professional standards” source. So it plausibly measured AI-written against human-written work, then labelled the AI bucket “watermarked”. No mechanism has been demonstrated by which a search engine reads a Claude statistical watermark. Expect the headline figure to be quoted for a year without the caveat its own author wrote.
What people actually search for
We pulled the keyword data for this piece rather than guessing at it. Two findings are worth publishing.
The first: “claude watermark remover” carries more monthly US search volume than “claude watermark” itself, 30 against 20 (Ahrefs, United States, captured 14 August 2026). More people want to strip the mark than to understand it, three days after it shipped. The same pattern shows up elsewhere. On Hacker News in May, a watermark-removal tool scored higher than the announcement of an industry watermarking standard published the same day.
The second: the questions people ask are not about generation. They are about accusation. “Are AI detectors reliable” runs at 800 US searches a month, “ai detection false positives” at 700, “can ai detectors be wrong” at 450, alongside a long tail of “falsely accused of using AI”, “student falsely accused of using AI”, and “what to do if you are falsely accused of using AI”. Nobody is searching for how to prove a machine wrote something. They are searching for how to prove one did not.
Six days later, the market was already selling a detector that cannot exist
We audited that demand on 17 August 2026, six days after the announcement, reading the pages in full rather than the search snippets. Four are worth publishing.
claudewatermark.com is titled “Claude Watermark Remover and Checker”. What it removes is zero-width spaces, joiners, soft hyphens, directional marks and em dashes. Anthropic’s text mark is none of those. Its own FAQ defines the thing it sells against as a pattern “that OpenAI may embed in ChatGPT-generated content”, and says that “while Anthropic has discussed the possibility of watermarking AI-generated text, the exact implementation and prevalence of watermarks in Claude outputs remains partially undisclosed” source. It sells a checker for a mark it calls hypothetical and attributes to the wrong company. Verisign’s registry record dates the domain to 10 March 2026, five months before the watermark existed source.
StealthGPT’s page is titled “Free Claude Watermark Remover and Detector” and offers “everything you need to catch, strip, and verify Claude’s mark”, including a “Detection Score on Every Response” so there is “no guessing whether a piece is actually clean”. The mechanism sits in the same paragraph: that score is “checked against the same tools that would flag Claude’s output downstream” source. Those tools are AI detectors. The product substitutes the category the FTC sanctioned for the reader that does not exist, then calls the result verification. It also promises to be “updated on an ongoing basis as Claude’s marking method changes”, for a method never published.
Human Writes advertises “Detect & Remove Claude & ChatGPT Watermarks”, strips invisible Unicode, rewrites the draft, and “then shows a built-in AI score” source. The same substitution.
Ninja Humanizer is the one telling the truth, and it is selling removal. Under a heading called “What this tool does not promise” it states that “Anthropic has not published its detector, and the mark is a statistical tilt rather than a stamp sitting in one spot. Anyone selling you a guaranteed removal rate is guessing.” It adds that “Turnitin, GPTZero and Originality are not looking for the Anthropic watermark. They score how predictable your writing is” source.
Two things follow. The most accurate public description of the detection situation we found anywhere sits on a page built to defeat the mark, not in the coverage of its launch. And the market’s working definition of Claude’s watermark has collapsed into em dashes and zero-width characters, because those are the only things anyone can point at. When a signal cannot be read, folk proxies fill the space.
None of this had to be built. Verisign’s records date ninjahumanizer.com to 20 November 2025 and humanwritesai.com to 4 November 2025, both months before the watermark shipped source. The evasion industry already existed at scale. The watermark gave it a new product page in under a week, and the same should be expected of every provenance feature that ships without a reader.
The regulatory shape is already settled. An advertised detection capability that cannot be substantiated is precisely what the FTC ordered Workado to stop, and we hold that order as a verified case file. Workado at least had a classifier that scored 53 percent against a measurable ground truth. For Claude’s watermark there is no public ground truth to score against at all.
What actually holds up, from our own files
We keep a registry of adjudicated cases in which AI-fabricated citations cost lawyers money, sanctions and licences. It is one of the larger such collections anywhere. Across all of it, not one court reached its finding by detecting AI.
They reached it by checking whether the citation existed.
United States v. Hayes is the cleanest statement of the principle. A federal magistrate judge found that a fictitious citation in a public defender’s brief bore “all the markings of a hallucinated case created by generative artificial intelligence (AI) tools such as ChatGPT and Google Bard”, then held she “need not make any finding” that AI had been used. The attorney denied using it. The sanction followed anyway, because the authority did not exist. Our case file on Hayes exists precisely because it marks where AI blame stops. The falsity was established. The authorship never was. Accountability did not require it.
Kohls v. Ellison closes the loop with an irony nobody would dare invent. Defending Minnesota’s political-deepfake law, the state filed an expert declaration from a Stanford AI-misinformation scholar. The declaration had been drafted with GPT-4o and not verified. It cited two articles that did not exist and mis-attributed a third. In January 2025 the court excluded it in its entirety. Expert evidence supporting a synthetic-media statute was itself unchecked synthetic output, and what caught it was not a detector. Someone looked up the citations.
That method is unglamorous, and it is what we apply to every file we publish: primary sources, named individuals on record, captures archived, the method printed on the page. Zalando’s 2018 restructuring, Klarna’s customer-service rollout, Moderna’s ChatGPT Enterprise deployment and Decagon’s work with Notion sit on this site because their claims were checked against records that exist independently of the companies making them. None of that required knowing which words a machine wrote.
A note on how this piece was researched
While assembling the regulatory sweep, our research kept surfacing confident, well-formatted claims about AI laws that do not exist. A UAE digital-economy and AI-content statute ratified by the Federal National Council. An Australian government response accepting mandatory AI guardrails in principle. Both came from aggregator sites that appear to be machine-generated. Both were plausible. Neither could be corroborated anywhere, so neither is in this piece.
No watermark would have caught that. Those sites were not passing off machine writing as human, and nobody cared who wrote them. They were publishing regulations that had never been enacted. What caught it was going to look for the law and finding nothing there.
Detection asks what made the content. Verification asks whether the claim inside it is true. A mark can tell you Claude was in the room. It cannot tell you whether the number in the third paragraph is real, and neither can a detector, a C2PA badge, or a label required by statute in four countries. That was never a detection problem. It was always a checking problem, and checking is work.
Sources
- Anthropic, “How Claude marks AI-generated content,” 11 August 2026. https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content
- TechCrunch, “Anthropic says it will watermark text generated by its AI models,” 11 August 2026. https://techcrunch.com/2026/08/11/anthropic-says-it-will-watermark-text-generated-by-its-ai-models/
- Kotaku, “YouTube Mistakenly Penalizes Popular Science Channel Kurzgesagt For ‘AI-Generated Slop’,” 8 August 2026. https://kotaku.com/youtube-mistakenly-penalizes-popular-science-channel-kurzgesagt-for-ai-generated-slop-2000722702
- PetaPixel, “Instagram Photos Are Being Labeled ‘Made With AI’ When They’re Not,” 28 May 2024. https://petapixel.com/2024/05/28/instagram-photos-are-being-labeled-made-with-ai-when-theyre-not/
- Digital Music News, “YouTube Confirms Altering Creators’ Videos Using AI,” 25 August 2025. https://www.digitalmusicnews.com/2025/08/25/youtube-alters-creator-videos-ai/
- Federal Trade Commission, “FTC Approves Final Order Against Workado, LLC,” August 2025. https://www.ftc.gov/news-events/news/press-releases/2025/08/ftc-approves-final-order-against-workado-llc-which-misrepresented-accuracy-its-artificial
- Sadasivan, Kumar, Balasubramanian, Wang, Feizi, “Can AI-Generated Text be Reliably Detected?” arXiv, March 2023. https://arxiv.org/abs/2303.11156
- Jovanovic, Staab, Vechev, “Watermark Stealing in Large Language Models,” ICML 2024. https://arxiv.org/abs/2402.19361
- Zhang, Edelman, Francati, Venturi, Ateniese, Barak, “Watermarks in the Sand: Impossibility of Strong Watermarking for Language Models,” ICML 2024. https://arxiv.org/abs/2311.04378
- Vanderbilt University, “Guidance on AI Detection and Why We’re Disabling Turnitin’s AI Detector,” 16 August 2023. https://www.vanderbilt.edu/brightspace/2023/08/16/guidance-on-ai-detection-and-why-were-disabling-turnitins-ai-detector/
- Inside Higher Ed, “Turnitin’s AI Detector: Higher-Than-Expected False Positives,” 1 June 2023. https://www.insidehighered.com/news/quick-takes/2023/06/01/turnitins-ai-detector-higher-expected-false-positives
- Liang, Yuksekgonul, Mao, Wu, Zou, “GPT detectors are biased against non-native English writers,” Patterns 4(7), July 2023. https://www.cell.com/patterns/fulltext/S2666-3899(23)00130-7
- Gizmodo, “AI Detectors Get It Wrong. Writers Are Being Fired Anyway,” 12 June 2024. https://gizmodo.com/ai-detectors-inaccurate-freelance-writers-fired-1851529820
- JCOM, “Visible sources and invisible risks: exploring the impact of AI disclosure on perceived credibility of AI-generated content,” 2026. https://jcom.sissa.it/article/pubid/JCOM_2501_2026_A09/
- Inside Higher Ed, “Adelphi Student Wins AI Plagiarism Lawsuit,” 11 February 2026. https://www.insidehighered.com/news/quick-takes/2026/02/11/adelphi-student-wins-ai-plagiarism-lawsuit
- Singapore Statutes Online, Online Safety (Relief and Accountability) Act 2025, section 16. https://sso.agc.gov.sg/Act/OSRAA2025?ProvIds=pr16-
- CNA, “AI-generated videos of Singapore politicians surged before GE2025,” 21 April 2025. https://www.channelnewsasia.com/singapore/ge2025-deepfakes-ai-generated-tiktok-politicians-disinformation-5078011
- EU AI Act, Article 50, applicable 2 August 2026. https://artificialintelligenceact.eu/article/50/
- Freshfields, “India targets deepfakes and AI-generated content: key changes under MeitY’s 2026 amendments to the IT Rules,” February 2026. https://www.freshfields.com/en/our-thinking/blogs/technology-quotient/india-targets-deepfakes-and-ai-generated-content-key-changes-under-meitys-2026-102mjwn
- Cyberspace Administration of China, Qinglang campaign on rectifying chaos in AI applications, 30 April 2026. https://www.cac.gov.cn/2026-04/30/c_1779289298718765.htm
- Google Search Central, “Spam policies for Google web search,” updated 15 May 2026. https://developers.google.com/search/docs/essentials/spam-policies
- Google, “New ways to identify AI-generated media online,” 19 May 2026. https://blog.google/innovation-and-ai/products/identifying-ai-generated-media-online/
- Ahrefs, “AI Overviews Cite AI-Generated Content More Than Human Writing,” 2026. https://ahrefs.com/blog/ai-overviews-cite-ai-generated-content-more-than-human-writing/
- Sean Goedecke, “Text AI watermarks will always be trivial to remove,” 2 July 2026. https://www.seangoedecke.com/text-ai-watermarks/
- First Page Sage, “Impact of AI Watermarking on SEO and GEO Results,” 11 August 2026. https://firstpagesage.com/seo-blog/impact-of-ai-watermarking-on-seo-geo-results/
- Dathathri et al., “Scalable watermarking for identifying large language model outputs,” Nature 634, 818-823, 23 October 2024. https://www.nature.com/articles/s41586-024-08025-4
- Google DeepMind, synthid-text reference implementation, Apache 2.0. https://github.com/google-deepmind/synthid-text
- Google DeepMind, SynthID product page, captured 17 August 2026. https://deepmind.google/technologies/synthid/
- Google, “SynthID Detector, a new portal to help identify AI-generated content,” 20 May 2025. https://blog.google/innovation-and-ai/products/google-synthid-ai-content-detector/
- OpenAI, “Understanding the source of what we see and hear online,” update of 4 August 2024. https://openai.com/index/understanding-the-source-of-what-we-see-and-hear-online/
- claudewatermark.com, “Claude Watermark Remover and Checker,” captured 17 August 2026. https://www.claudewatermark.com/
- StealthGPT, “Free Claude Watermark Remover and Detector,” captured 17 August 2026. https://www.stealthgpt.ai/use-cases/claude-watermark-remover
- Human Writes, “AI Watermark Remover,” captured 17 August 2026. https://humanwritesai.com/watermark-checker
- Ninja Humanizer, “Claude Watermark Remover,” captured 17 August 2026. https://ninjahumanizer.com/claude-watermark-remover
- Verisign, Registry Data Access Protocol records for claudewatermark.com, ninjahumanizer.com and humanwritesai.com, retrieved 17 August 2026. https://rdap.verisign.com/com/v1/domain/claudewatermark.com
Keyword volumes cited in “What people actually search for” are from Ahrefs Keywords Explorer, United States, captured 14 August 2026. Sources 31 to 34 are vendor marketing pages, quoted as evidence of what is being advertised, not as technical authority. Domain registration dates are from the .com registry’s own RDAP service.
Questions
Does Claude watermark text, and what does the Claude watermark actually prove?
Yes. Claude models launched on or after 2 August 2026 embed a statistical watermark in generated text, with C2PA metadata on generated image files. Anthropic's own documentation says a detected mark indicates content may have been processed by Claude and is not fully conclusive, and notes that people use Claude to proofread, translate and summarise their own writing. Proofreading leaves the same mark as full generation.
Are AI detectors reliable?
The measured record says no. The FTC ordered Workado to stop advertising its AI content detector as 98 percent accurate after independent testing put it at 53 percent on general-purpose content. Turnitin conceded a higher-than-expected false-positive rate within weeks of launch, and Vanderbilt disabled it after calculating that roughly 750 of its 75,000 papers could have been wrongly flagged.
How common are AI detection false positives, and who do they hit hardest?
They fall disproportionately on non-native English speakers. In Patterns, seven detectors misclassified more than half of TOEFL essays written by non-native speakers as AI-generated while performing near-perfectly on US eighth-grade essays. The researchers confirmed the cause by enriching the vocabulary of the non-native essays, which made the flags disappear.
Can an AI watermark be removed or faked?
Both. Anthropic lists heavy editing, paraphrasing and translation as things that weaken or destroy the mark. Separately, ICML research on watermark stealing showed an attacker can scrub and forge state-of-the-art schemes for under fifty dollars at over 80 percent success, which means a human's writing can be stamped as machine-made.
Can anyone outside Anthropic check whether text carries the Claude watermark?
Not as of 17 August 2026. Anthropic says it is working to enable users and third parties to detect the mark and will publish detection details later. No lab has shipped a public text read path: Google open-sourced its SynthID-Text detector, but detection needs the watermarking key, and Google's own verification portal covers image, video and audio, not text. OpenAI built a text watermark and declined to release it, citing trivial circumvention and the risk of stigmatising non-native English speakers.
Do search engines rank AI-generated content lower for carrying a watermark?
No published policy says so. Google's spam policy triggers on scaled content abuse, meaning many pages generated to manipulate rankings without adding value, never on whether AI was used. Google has placed provenance verification under Search appearance rather than ranking, and has never stated that watermarks affect ranking in either direction.
Sources
- Anthropic, How Claude marks AI-generated content , 2026-08-11
- TechCrunch, Anthropic says it will watermark text generated by its AI models , 2026-08-11
- Kotaku, YouTube Mistakenly Penalizes Popular Science Channel Kurzgesagt For 'AI-Generated Slop' , 2026-08-08
- PetaPixel, Instagram Photos Are Being Labeled 'Made With AI' When They're Not , 2024-05-28
- Digital Music News, YouTube Confirms Altering Creators' Videos Using AI , 2025-08-25
- Federal Trade Commission, FTC Approves Final Order Against Workado, LLC, Which Misrepresented the Accuracy of its AI Content Detection Product , 2025-08
- Sadasivan, Kumar, Balasubramanian, Wang, Feizi (arXiv / TMLR), Can AI-Generated Text be Reliably Detected? , 2023-03
- Jovanovic, Staab, Vechev (ICML 2024), Watermark Stealing in Large Language Models , 2024-02
- Zhang, Edelman, Francati, Venturi, Ateniese, Barak (ICML 2024), Watermarks in the Sand: Impossibility of Strong Watermarking for Language Models , 2023-11
- Vanderbilt University, Guidance on AI Detection and Why We're Disabling Turnitin's AI Detector , 2023-08-16
- Inside Higher Ed, Turnitin's AI Detector: Higher-Than-Expected False Positives , 2023-06-01
- Liang, Yuksekgonul, Mao, Wu, Zou. Patterns 4(7), GPT detectors are biased against non-native English writers , 2023-07
- Gizmodo, AI Detectors Get It Wrong. Writers Are Being Fired Anyway , 2024-06-12
- JCOM (Journal of Science Communication), Visible sources and invisible risks: exploring the impact of AI disclosure on perceived credibility of AI-generated content , 2026
- Inside Higher Ed, Adelphi Student Wins AI Plagiarism Lawsuit , 2026-02-11
- Singapore Statutes Online, Online Safety (Relief and Accountability) Act 2025, section 16, inauthentic material abuse , 2025-11-25
- CNA, AI-generated videos of Singapore politicians surged before GE2025 , 2025-04-21
- EU AI Act, Article 50: Transparency Obligations for Providers and Deployers of Certain AI Systems , 2026-08-02
- Freshfields, India targets deepfakes and AI-generated content: key changes under MeitY's 2026 amendments to the IT Rules , 2026-02
- Cyberspace Administration of China, Qinglang campaign on rectifying chaos in AI applications , 2026-04-30
- Google Search Central, Spam policies for Google web search , 2026-05-15
- Google, New ways to identify AI-generated media online , 2026-05-19
- Ahrefs, AI Overviews Cite AI-Generated Content More Than Human Writing , 2026
- Sean Goedecke, Text AI watermarks will always be trivial to remove , 2026-07-02
- First Page Sage, Impact of AI Watermarking on SEO and GEO Results , 2026-08-11
- Dathathri et al., Nature 634, 818-823, Scalable watermarking for identifying large language model outputs , 2024-10-23
- Google DeepMind, synthid-text, reference implementation of SynthID-Text watermarking and detection (Apache 2.0) , 2024-10
- Google DeepMind, SynthID, product page stating the Detector portal verifies image, video or audio , 2026-08-17
- Google, SynthID Detector, a new portal to help identify AI-generated content , 2025-05-20
- OpenAI, Understanding the source of what we see and hear online, update of 4 August 2024 on text watermarking , 2024-08-04
- claudewatermark.com, captured 17 August 2026, Claude Watermark Remover and Checker , 2026-08-17
- StealthGPT, captured 17 August 2026, Free Claude Watermark Remover and Detector , 2026-08-17
- Human Writes, captured 17 August 2026, AI Watermark Remover, Detect & Remove Claude & ChatGPT Watermarks , 2026-08-17
- Ninja Humanizer, captured 17 August 2026, Claude Watermark Remover: Free Tool to Clean Claude Text , 2026-08-17
- Verisign, retrieved 17 August 2026, Registry Data Access Protocol record, claudewatermark.com , 2026-08-17
This is analysis, not a verified outcome. It carries no verification badge and never will. The proof lives in the case files, where every figure is checked against the public record and the method is printed on the page.