OpenAI brings text watermarking to its API — and unlike Anthropic, it’s off by default
The New Stack Paul Sawers ● Covered by 5 sources
OpenAI is adding text watermarking to its API, but only if developers switch it on. It’ll auto-tag some ChatGPT and Codex text in the EU, while Anthropic went global by default.
Based on reporting by The New Stack, Paul Sawers — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
OpenAI is letting developers opt in to text watermarking for API output, bringing its provenance push to a harder problem: plain language. The company says the new system, textGrain, works by nudging model word choice just enough to leave a statistical pattern behind. Over longer passages, that pattern becomes something its detector can spot.
The rollout starts today for API customers worldwide on supported models, but the switch is off by default. OpenAI says that matters because customers can decide how watermarking fits their own transparency obligations. In the European Union, the company plans to go further in the coming weeks and automatically watermark eligible text from ChatGPT and Codex to meet EU AI Act transparency requirements.
That puts OpenAI on a different path from Anthropic. When Anthropic announced watermarking for Claude in August, it said the feature would apply globally because it didn’t have a reliable regional limiter. OpenAI, by contrast, is giving API customers control at the project or organization level, with no extra request changes needed once watermarking is turned on.
OpenAI is also making a familiar argument for building its own system. The company says textGrain gives it more control over the tradeoff between detectability and how varied the model’s answers remain. It says the method matched or beat SynthID in testing, and it plans to open-source the technology so others can build on it.
The catch is that text watermarking is still a fragile thing. OpenAI says detection is around 80% for watermarked 200-token passages and 95% for 400-token passages in areas like psychology, at a 1% false-positive target. But the signal drops in more constrained writing such as math, and editing can wreck it fast: swapping out 10% of the words in a 400-token passage cut detection from around 92% to 66%, while 25% replacement dropped it to 17%.
My take — AI-written commentary, not fact-checked reporting
OpenAI is doing the sensible thing here: give users a switch, don’t pretend every output needs the same treatment, and don’t make the API feel like a surveillance product. The awkward part is that the company still wants the aura of provenance without admitting how easily a determined user can sand it off. Watermarking is useful, but only if people stop treating it like a seatbelt for the whole car.
Read more about this at: The New Stack