TLDRocket
Sign in

AI Models Are Watermarking Text—Will You Notice?

IEEE Spectrum Matthew S. Smith

Anthropic says all future Claude models will watermark their text. That could help spot AI, but it may also nudge the writing itself.

Based on reporting by IEEE Spectrum, Matthew S. Smith — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Anthropic has started putting text watermarks on the road map. On 11 August, the company said all future Claude models will generate output that contains a watermark identifying it as AI-made. Google already uses its own text watermark on Gemini, and Anthropic says its version is based on that system. OpenAI, for now, still hasn’t shipped one, though it says it plans to.

The timing is not accidental. The European Union’s AI Act requires watermarks for AI models released after 2 August, 2026, alongside other rules meant to curb deceptive or manipulative AI content. The act also calls for watermarks on images, audio, and video. Those media watermarks are already common, with detection rates above 99 percent in some cases. Text is the hard part.

That’s because a text watermark is not a hidden tag or some invisible character trick. It works by quietly changing how a model chooses words. Large language models assign probabilities to possible next tokens, then sample from that distribution. The classic 2023 approach from John Kirchenbauer and colleagues splits words into red and green lists, then nudges the green set to appear a bit more often. Their paper reported a 98.4 percent detection rate and zero false positives in responses of about 200 tokens.

But the same subtlety that makes the watermark hard to see also makes people nervous. John Gruber calls it a “perversion of writing” and argues that the text is no longer quite what the model would have produced. Kirchenbauer takes the opposite view: if the watermark is detectable, then of course it changed something, but the real question is whether that change hurts usefulness. That argument gets messy fast when replies are short. Vinu Sankar Sadasivan of Meta says a 20-word tweet may need 50 or 60 percent of its words drawn from the green list for the watermark to show up clearly.

Google’s own 2024 work on SynthID-Text found no significant difference in user feedback across 20 million Gemini responses when comparing watermarked and unwatermarked output. Even so, the company’s paper showed detection rates that can drop below 50 percent on short replies. Anthropic hasn’t explained how it is balancing watermark strength against output quality, and it declined to add more detail for this story.

And the story is getting bigger than labeling. Researchers are now looking at text watermarks as a way to track where data goes. A 2026 paper co-authored by Kirchenbauer says models trained on watermarked text can themselves produce watermarked output, which could help trace training data or keep models from feeding on their own recycled mistakes.

My take — AI-written commentary, not fact-checked reporting

This is the sort of regulation-by-footnote that EU tech policy loves: noble goal, fuzzy mechanics, and a lot of confidence about tradeoffs other people will pay for. Watermarks may be useful for provenance, but pretending they’re free and invisible is the usual Silicon Valley bedtime story. The real test is whether vendors admit the quality cost when the prompt is short and the stakes are high.

Read more about this at: IEEE Spectrum

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.