Digital
OpenAI to roll out invisible text watermarks for ChatGPT and Codex in EU
API customers can opt in globally as OpenAI warns text detection still has limits
MUMBAI: OpenAI is putting a digital fingerprint on AI-generated text, announcing plans to introduce invisible text watermarks for eligible ChatGPT and Codex outputs in the European Union in response to the EU AI Act.
The company said the rollout will take place over the coming weeks and will initially be limited to the EU. At the same time, OpenAI will allow API customers worldwide to opt in to text watermarking for select models, although the feature will remain switched off by default.
The move comes as the EU AI Act requires providers of generative AI systems to make AI-generated content identifiable in a machine-readable format. OpenAI, however, stressed that text watermarking and detection remain relatively young technologies and come with significant limitations.
OpenAI’s system, called textGrain, adds an invisible statistical signal to a model’s choice of words. A dedicated detector can then analyse a passage to determine whether it contains an OpenAI watermark.
The company said its evaluations found textGrain matched or exceeded the performance of other approaches it tested, including Google’s SynthID for text. However, OpenAI cautioned that strong results under controlled conditions do not necessarily translate into reliable detection in everyday use.
The technology faces two basic challenges: false positives, where a watermark is detected despite not being present, and false negatives, where an existing watermark is missed.
Detection also varies significantly according to the length and type of text. OpenAI said that, at a target false-positive rate of 1 per cent, its detector identified watermarks in around 80 per cent of 200-token psychology passages, rising to about 95 per cent for 400-token passages.
Performance was considerably weaker for subjects such as mathematics, where there is less flexibility in word choice.
Editing can also knock the watermark off course. In tests involving 400-token passages, replacing 10 per cent of words with synonyms reduced detection from around 92 per cent to 66 per cent. Replacing 25 per cent of the words brought detection down to 17 per cent.
OpenAI said it plans to make the underlying watermarking technology available as open source so that others can build on it.
Despite the watermark being embedded in the model’s word choices, OpenAI said its tests showed no meaningful performance difference between watermarked and unwatermarked outputs from its latest frontier model, Astra.
Across benchmarks including Artificial Analysis Intelligence Index, AutomationBench, DeepSWE, Terminal-Bench, BrowseComp, HealthBench Professional and GPQA Diamond, the results for watermarked and unwatermarked outputs remained broadly comparable.
The company said this suggests watermarking can be introduced without materially affecting the quality of model responses.
OpenAI is also taking care to put limits around what its detector can actually prove.
A detected watermark does not measure how much a human contributed to a passage. It can indicate that an OpenAI system generated or processed some of the text, but cannot determine the extent of human judgement, editing or creativity involved.
Nor does it establish ownership, legal responsibility or whether disclosure of AI use was required.
The signal does not identify the user, organisation, account, prompt or conversation associated with a piece of text. It also cannot verify whether the content itself is accurate, misleading, harmful or presented in the right context.
Conversely, a failure to detect a watermark does not prove that a human wrote the text. OpenAI said short passages, edited or translated content, unsupported models and older outputs may all evade detection.
OpenAI is opening applications for access to its text watermark detector, but will initially limit access to approved researchers and expert organisations.
The approach is intended to help evaluate the technology and develop responsible use cases before wider access is considered. The detector will report whether an OpenAI watermark is detected, without identifying users or exposing their prompts or conversations.
The company said it expects to review this policy as evidence, standards and technology evolve.
For API users, meanwhile, watermarking is being made available as an opt-in feature globally for select models. OpenAI is also working with cloud partners to extend watermarking to OpenAI model outputs accessed through their services.
OpenAI’s wider content provenance strategy already covers images and audio through technologies including Content Credentials, C2PA standards and invisible SynthID watermarks. It also offers verification tools for image and audio content.
Text presents a tougher challenge because it can be rewritten, translated and edited with relative ease, potentially weakening or removing a watermark signal.
OpenAI said it will continue studying how its watermarking technology performs after editing and translation, while exploring ways to distinguish AI assistance from AI authorship more meaningfully.
For now, the company is taking a cautious route: watermark first, learn from real-world use and widen detection access only when the results can be interpreted responsibly.




