TLDRocket
Sign in

OpenAI brings text watermarking to its API — and unlike Anthropic, it’s off by default

The New Stack Paul Sawers ● Covered by 5 sources

OpenAI is adding text watermarking to its API, but only if developers switch it on. It’ll auto-tag some ChatGPT and Codex text in the EU, while Anthropic went global by default.

Based on reporting by The New Stack, Paul Sawers — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

OpenAI is letting developers opt in to text watermarking for API output, bringing its provenance push to a harder problem: plain language. The company says the new system, textGrain, works by nudging model word choice just enough to leave a statistical pattern behind. Over longer passages, that pattern becomes something its detector can spot.

The rollout starts today for API customers worldwide on supported models, but the switch is off by default. OpenAI says that matters because customers can decide how watermarking fits their own transparency obligations. In the European Union, the company plans to go further in the coming weeks and automatically watermark eligible text from ChatGPT and Codex to meet EU AI Act transparency requirements.

That puts OpenAI on a different path from Anthropic. When Anthropic announced watermarking for Claude in August, it said the feature would apply globally because it didn’t have a reliable regional limiter. OpenAI, by contrast, is giving API customers control at the project or organization level, with no extra request changes needed once watermarking is turned on.

OpenAI is also making a familiar argument for building its own system. The company says textGrain gives it more control over the tradeoff between detectability and how varied the model’s answers remain. It says the method matched or beat SynthID in testing, and it plans to open-source the technology so others can build on it.

The catch is that text watermarking is still a fragile thing. OpenAI says detection is around 80% for watermarked 200-token passages and 95% for 400-token passages in areas like psychology, at a 1% false-positive target. But the signal drops in more constrained writing such as math, and editing can wreck it fast: swapping out 10% of the words in a 400-token passage cut detection from around 92% to 66%, while 25% replacement dropped it to 17%.

My take — AI-written commentary, not fact-checked reporting

OpenAI is doing the sensible thing here: give users a switch, don’t pretend every output needs the same treatment, and don’t make the API feel like a surveillance product. The awkward part is that the company still wants the aura of provenance without admitting how easily a determined user can sand it off. Watermarking is useful, but only if people stop treating it like a seatbelt for the whole car.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.