OpenAI’s textGrain Watermark Is Coming to EU ChatGPT: What It Can Detect

OpenAI will watermark eligible EU ChatGPT and Codex text with textGrain. Its own tests show why detection cannot prove authorship.
European Union flags outside the Berlaymont building in Brussels
EU flags outside the European Commission’s Berlaymont headquarters in Brussels. Photo: Hlynz / Wikimedia Commons, CC BY-SA 3.0.

OpenAI will begin adding an invisible text watermark to eligible ChatGPT and Codex responses in the European Union over the coming weeks, extending a regulatory requirement into two of its most widely used writing and coding products.

The system, called textGrain, is already available as an opt-in setting for API customers worldwide. It changes the statistical pattern of the model’s word choices rather than inserting hidden characters, spaces or metadata. OpenAI plans to make the technique open source, but access to its detector will initially be restricted to approved researchers and expert organizations.

That restriction matters because textGrain is a provenance signal, not a reliable test of who wrote a document. OpenAI’s own October 5 announcement reports that modest editing can sharply weaken detection, while short answers, code and tightly constrained factual writing may not carry enough signal to detect at all.

How textGrain changes a model’s word choices

A language model writes one token at a time by assigning probabilities to possible next words or word fragments. textGrain divides those candidates into key-dependent groups and subtly changes which group the sampler favors. Within the selected group, the original relative probabilities remain intact.

The result is a pattern that a detector holding the same secret key can reconstruct from the finished text. It does not need the original prompt, conversation, generating model or watermark-strength setting. The detector scores whether observed token choices align with the keyed pattern more often than chance would predict.

OpenAI’s technical report describes a more specific mechanism than the familiar shorthand of choosing from “green” and “red” token lists. textGrain uses optimal transport to couple token groups with keyed randomness, then limits how much sampling entropy the watermark may remove. In plainer terms, the system is designed to leave a detectable preference without making repeated answers rigid or obviously changing their wording.

This differs from adding a tag to a file. Copying and pasting watermarked prose preserves its wording and may preserve the signal; exporting it to plain text does not automatically strip anything. Rewriting the passage can.

OpenAI’s numbers show the practical limits

At a target false-positive rate of 1%, OpenAI detected the watermark in about 80% of 200-token psychology answers and roughly 95% of 400-token answers. Performance was substantially lower for mathematics, where the model has fewer equally valid ways to express the answer.

Editing had a larger effect. In tests on 400-token English passages, replacing 10% of the words with synonyms reduced detection from about 92% to 66%. Replacing 25% cut it to 17%.

Those figures create several important boundaries:

  • A negative result does not prove human authorship. The text may be too short, translated, edited, generated by an unsupported or older model, or produced by another AI service.
  • A positive result does not measure human contribution. The watermark may reflect generation, substantial editing or processing by an OpenAI model. It cannot say how much judgment or original work came from a person.
  • The detector does not identify an account. OpenAI says the watermark does not encode a user, organization, prompt or conversation.
  • It does not verify truth or ownership. Detection says nothing about factual accuracy, copyright, permission or legal responsibility.

Short passages and source code present a structural problem, not merely a model-tuning problem. A watermark needs repeated points where several next-token choices are acceptable. Exact answers, quotations, copied source text and executable code offer fewer such choices. OpenAI’s guidance says the EU transparency code does not require watermarking for outputs under 200 tokens, roughly 150 English words, or for code snippets.

The detector will not be a public AI-writing checker at launch

OpenAI is accepting applications from researchers and expert organizations, but it is not giving every teacher, editor, employer or website operator a public text checker. The company cites missed watermarks and false positives as reasons for the controlled rollout. Its publicly available verification tools will continue to cover supported images and audio, not text.

That makes textGrain materially different from the many commercial “AI detector” products that classify prose by style. A watermark detector checks for a deliberately embedded keyed pattern. A classifier guesses from features associated with model-written text and may label unwatermarked writing. Neither result, by itself, establishes misconduct.

Anthropic has taken a related but broader deployment path. Its Claude watermark uses Google’s SynthID-Text method, applies globally to newer models, and has a detection API in private preview for eligible organizations. OpenAI is limiting the automatic textGrain rollout to eligible ChatGPT and Codex users in the EU at launch, while leaving API watermarking optional worldwide.

The vendors share important caveats: longer, varied prose gives a detector more evidence; factual passages and code give it less; editing can erode the signal; and a detected watermark shows likely model involvement rather than authorship.

Why the EU AI Act is driving the rollout

Article 50 of the EU AI Act requires providers of generative systems to mark AI-generated or manipulated output in a machine-readable, detectable form, using techniques that are effective and reliable as far as technically feasible. The European Commission’s transparency code covers audio, images, video and text while explicitly accounting for technical limitations and implementation cost.

The law separates provider marking from deployer disclosure. Publishers of AI-generated or manipulated text about matters of public interest may need a visible disclosure unless the material has undergone human review and carries editorial responsibility. An invisible watermark therefore does not automatically replace labels, editorial review or an organization’s own record keeping.

For API developers, OpenAI allows watermarking to be enabled at the organization or project level for supported models. That setting does not grant detector access. Teams also need to check whether their cloud provider supports the signal, because availability through distribution partners will roll out separately.

A sensible policy for schools, employers and publishers

Organizations likely to receive detector access should define how results will be used before checking real work. A defensible process should include:

  1. Treat detection as one provenance clue. Do not convert a probabilistic result into an automatic cheating, hiring or disciplinary decision.
  2. Preserve context. Keep drafts, version history, citations, prompts disclosed under policy and the original file where appropriate. Those records can show a writing process that a watermark cannot.
  3. Account for length and genre. A short answer, code sample, translation or factual response is a poor candidate for a strong conclusion.
  4. Provide human review and an appeal path. People should be able to explain permitted AI assistance, editing workflows and false or incomplete signals.
  5. Separate authorship from accuracy. Fact-check a passage regardless of whether a watermark is present.

Developers enabling textGrain should also test their real production output rather than relying only on general benchmarks. Customer-support macros, legal clauses, code-generation tools and multilingual products all constrain language differently. Detection rates measured on open-ended English answers may not predict performance in those settings.

What remains unanswered

OpenAI has not yet published a complete supported-model list in the announcement; customers see current options in their API settings. Cloud-partner timing, the scope of “eligible” ChatGPT and Codex output, and the schedule for broader detector access also remain open.

The company says its benchmark results show no meaningful quality loss on Astra, but those are internal evaluations. Independent researchers will need detector access, implementation details and representative samples to test false-positive rates across languages, genres and adversarial edits.

textGrain gives platforms and investigators a stronger signal than stylistic guesswork when the watermark survives. OpenAI’s measurements also make the limit unusually clear: the signal weakens quickly as text is shortened or rewritten. The responsible use of the detector starts with refusing to ask it a question it cannot answer: who, exactly, wrote this?

Leave a Reply

Your email address will not be published. Required fields are marked *

Previous Post
Google promotional artwork for the Gemini 4 Argon frontier AI model

Gemini 4 Argon Debuts With 1M-Token Output, but Access Starts With Cyber Defenders

Next Post
Model Context Protocol logo and protocol illustration

MCP Protocol Pivoting Exposes the Trust Gap Between AI Agents

Related Posts