TLDRocket
Sign in

Claude Fable 5.1 watermark: It has a blind spot developers can’t ignore

The New Stack Amanda Caswell Covered by 2 sources

Anthropic put a watermark in Claude Fable 5.1’s text. It’s much weaker on code, which is the part developers may care about most.

Based on reporting by The New Stack, Amanda Caswell — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Anthropic rolled out Claude Fable 5.1 on Tuesday with a built-in statistical watermark meant to help show when the model has been involved in writing text. The catch is right there in the fine print: the signal is far easier to find in natural language than in code, and code is exactly where a lot of developers will care most.

The system doesn’t tack on hidden characters or metadata. It tweaks the randomness Claude uses when it picks between possible next tokens, and Anthropic says that shouldn’t change the quality or meaning of the output. Over a long enough passage, those choices leave a detectable pattern. But code is a much tighter space. Pick the wrong variable, operator, or function and the program behaves differently, or stops working altogether.

So Anthropic backs off when precision matters. If a token is required for accuracy, the watermark isn’t applied. That means the signal may show up in comments or other looser parts of a response, while short code outputs may not carry enough of it to detect reliably. Anthropic is also opening detection through an API in private preview, but only for selected groups for now: regulators, law enforcement, media organizations, fact-checkers, researchers, and enterprises that need it for AI Act compliance.

The company says it will apply the watermark worldwide and plans to extend it to older Claude models over the coming months. It’s based on Google DeepMind’s SynthID-Text, and because it lives in the text itself, copying the output elsewhere won’t strip it away. Anthropic says some editing can survive too, though enough rewriting will wipe it out.

Fable 5.1 also tightens a separate issue around preserved thinking blocks in the Messages API. Those encrypted blocks can carry reasoning forward across turns, but Anthropic says that if developers change earlier conversation context while keeping the blocks, Claude can decrypt and print its reasoning, which could then be used to train another model. Starting with new accounts created on or after Aug. 31, that route is being closed across Claude Platform, Amazon Bedrock, Google Cloud Vertex AI and Microsoft Azure Foundry. Existing accounts can keep using Fable 5.1 without the restriction for now, but Anthropic says future model releases will apply it to everyone.

My take — AI-written commentary, not fact-checked reporting

This is Anthropic doing the classic AI-company thing: promising transparency while quietly making the useful bits harder to use. The watermark sounds noble until it hits code, where reality still beats policy. And the thinking-block restriction is the bigger tell anyway — the industry keeps calling these systems flexible, then spends half its time locking down the ways people actually use them.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.