OpenAI has outlined a staged approach to deploying text watermarking as part of its compliance strategy for the EU AI Act. The company described a system called textGrain that embeds an invisible statistical signal in model-generated word choices, and a matching detector that searches for that signal to indicate whether an OpenAI system generated or processed a passage.
OpenAI said it will publish and update a technical report on textGrain and intends to release the work as open source so others can build on the approach. The company framed text watermarking as one element of a layered provenance strategy that also includes open standards, Content Credentials, SynthID watermarks for supported images and audio, and verification APIs such as its web Verify tool and the Content Provenance API.
The watermarking approach is not presented as foolproof. OpenAI reported experimental results showing the detector’s performance varies significantly with passage length, the domain of the text, and post-generation editing. Using a 1% target false positive rate, the detector identified watermarks in roughly 80% of 200-token passages and about 95% of 400-token passages for topics with flexible wording, such as psychology. Detection rates were substantially lower on material with constrained wording, such as mathematics.
Editing reduces detectability. In tests on 400-token passages, substituting 10% of words with synonyms dropped detection from about 92% to 66%, and substituting 25% of words reduced detection to 17%. Based on those limitations and the risks of both false positives (detecting a watermark where none exists) and false negatives (failing to detect an existing watermark), OpenAI said it will restrict early access to the detector to vetted researchers and expert organizations rather than releasing it publicly at launch.
OpenAI also examined how watermarking affects model outputs and benchmarks. The company reported no meaningful differences across the benchmark suite it used to evaluate Astra, a frontier model it tested, with scores largely comparable with and without watermarking across the listed tests. That suggests limited observable impact on evaluated performance under the conditions the company examined, though OpenAI cautioned that controlled results do not guarantee reliable detection in everyday use.
For deployment, OpenAI described a phased rollout. API customers worldwide can opt in to receive watermarked text outputs for select models starting immediately; watermarking will remain off by default in the API. Over the coming weeks, the company plans to add an invisible watermark to eligible ChatGPT and Codex text outputs for users in the European Union across plans. OpenAI said it is not enabling watermarking as a global default at launch.
Access to the watermark detector will be available by application. Approved researchers and expert organizations can apply for detector access to help evaluate reliability and responsible uses. OpenAI said the detector reports whether it finds an OpenAI watermark without identifying users, prompts, or conversations. The company emphasized it will not make the detector publicly available at launch because missed watermarks and false positives could produce harms if the tool were used irresponsibly.
OpenAI was explicit about what a detected watermark does and does not mean. A detected watermark may indicate OpenAI generated or processed part of a passage, but it does not measure the extent of human contribution, establish ownership or legal responsibility, identify the user, or verify factual accuracy. Conversely, absence of a detected watermark does not prove human authorship: text may be too short, too heavily edited or translated, produced by an unsupported model, generated before watermarking was applied, or created by another provider’s tools.
The company committed to continuing research into watermark robustness under editing and translation, improving detection, and revising its approach as standards, evidence, and regulatory requirements evolve. OpenAI said it will expand detector access only when it believes results can be interpreted responsibly, and highlighted provenance signals as part of a broader effort that combines policy, abuse detection, reporting, and enforcement to help platforms and people assess potentially misleading AI-generated content.
Taken together, OpenAI’s plan positions text watermarking as one of several provenance tools intended to increase transparency about AI-generated content while acknowledging significant technical limits. The phased rollout, limited initial detector access, and ongoing research emphasize a cautious approach as the company adapts to regulatory requirements and seeks to improve watermarking reliability.
Source: Read the original source

Leave a Reply