TLDRocket
Sign in

Anthropic shares more details about how Claude’s new watermarks will work

TechCrunch Anthony Ha Covered by 2 sources

Anthropic explained how Claude’s new text watermarks will work. It says normal readers won’t notice, but edits and code limit how much can be tagged.

Based on reporting by TechCrunch, Anthony Ha — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Anthropic is trying to cool off the backlash around Claude’s new text watermarking by explaining, in plain terms, what the system does and doesn’t do. The company published a blog post on Friday after users started arguing about the change earlier this week, when Anthropic said it would add watermarking to meet the EU AI Act’s Transparency Code.

The basic idea is simple enough. When Claude is making low-stakes choices — Anthropic used the example of picking between “overcast” and “grey” — it can build a pattern into the response that people won’t see, but a detector with the right key can. Anthropic says this won’t change output quality, and that a watermarked response should look identical to an unwatermarked one.

Under the hood, Anthropic says it will use the SynthID-Text approach described by Google DeepMind in 2024, and it plans to release a watermark detection API. It also drew a line between watermarking and the sort of AI-detection tools sold by companies like Pangram, which look for stylistic tells in writing. Those are not the same thing, Anthropic said, because spotting writing patterns is different from checking for a watermark.

The company also tackled the obvious workaround: just rewrite the text. Light edits probably won’t erase the watermark fully, Anthropic said, but a full rewrite will. And if Claude has only proofread or lightly edited something, there may be very little watermark left to detect in the first place because nearly all the words came from the human author.

Code is a messier case. Anthropic says code should carry less of a watermark than ordinary text because the model has to produce working code, which leaves less room for arbitrary word choice. Comments can still carry some watermarking, but the company says the effect on the actual code should be negligible. Anthropic also made clear this won’t be a Claude-only feature; it says other major model developers signed the same Code of Practice and will build their own watermarks too.

My take — AI-written commentary, not fact-checked reporting

This is the part of AI regulation that actually makes sense: invisible, specific, and aimed at disclosure instead of theater. The interesting bit is that Anthropic is basically admitting the watermark matters most where the model has wiggle room, which is also where people tend to get nervous. The internet can rage all it wants, but compliance work is now shipping, and the rest of the industry is coming along for the ride.

Read more about this at: TechCrunch

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.