AI content provenance & transparency

OpenAI text watermarking: what textGrain can and cannot prove

•Make Better Editorial

OpenAI is rolling out invisible text watermarks in the EU and optional API watermarking globally. Here’s what detection can signal—and why it cannot prove authorship, accuracy or ownership.

OpenAI is beginning a phased rollout of text watermarking in response to the EU AI Act. API customers worldwide can now opt in to watermarked output for select models, while eligible ChatGPT and Codex text in the European Union is scheduled to receive an invisible watermark over the coming weeks. The change creates a new provenance signal for AI-generated text—but OpenAI’s own results show why teams should not treat detection as a binary authorship test.

Rollout boundary

API watermarking is opt-in and off by default. The planned ChatGPT and Codex rollout applies to eligible text output in the EU, not as a global default at launch. Detector access is initially limited to approved researchers and expert organizations.

How textGrain works

OpenAI calls its system textGrain. It adds an invisible statistical signal through the model’s word choices, and a detector looks for that signal later. OpenAI says it plans to open-source the technology, but also emphasizes that performance measured under controlled conditions does not guarantee reliable detection in everyday use.

Detection gets weaker with less text and more editing

~80% detected
200-token passage
Psychology content at a 1% target false-positive rate
~95% detected
400-token passage
Same evaluation setting
92% → 66%
10% synonym replacement
400-token English passages
92% → 17%
25% synonym replacement
Same editing evaluation
Make Better analysis

For marketers, publishers and AI teams, the important operational lesson is that provenance and authorship are different questions. A watermark can be useful as one machine-readable signal in a content pipeline, but a missing signal should not be converted into “human-written,” and a detected signal should not be converted into “fully AI-authored.” Editing, translation, short passages and unsupported models can all weaken that inference.

What a watermark does not prove

Signal vs conclusion

A watermark can help signalIt does not establish
An OpenAI system generated or processed some textHow much human judgment or editing was involved
A statistical provenance marker is detectableWho created the text or which account or prompt was used
The passage carries an OpenAI-origin signalOwnership, legality or responsibility
A detector found its expected statistical patternWhether the text is accurate, safe or trustworthy
No watermark was detectedThat the text was written by a human

Why the EU context matters

The European Commission says Article 50 transparency obligations apply from August 2, 2026. Its Code of Practice covers machine-readable marking and detection for providers, while deployer obligations include disclosure for certain AI-generated or manipulated public-interest text. The Commission also notes an exception where public-interest text has undergone human review and is subject to editorial responsibility. Teams should treat the regulation and the watermarking technology as related but separate layers: a technical signal does not by itself decide whether a specific disclosure obligation applies.

A practical provenance workflow

  1. Record which model and workflow produced or transformed content at generation time instead of relying only on later detection.
  2. If you use the OpenAI API and need machine-readable provenance, test the opt-in watermark on representative content before making it part of a compliance workflow.
  3. Keep human-review and editorial-responsibility records separate from watermark status; the watermark does not measure human contribution.
  4. Treat detector results as probabilistic evidence. Define what happens on a positive, negative or uncertain result before using it for moderation or publishing decisions.
  5. Retest after common edits, translation and shortening because those transformations can materially weaken detection.
  6. For regulated publishing decisions, map the workflow to the applicable EU requirements rather than assuming the watermark itself satisfies every disclosure duty.

The strongest reason not to over-read this launch

Text watermarking is fragile in exactly the situations common in real publishing: short copy, rewriting, translation and constrained language. OpenAI is limiting detector access at launch because false positives and missed watermarks still matter. That makes textGrain useful as an additional provenance layer, not a universal AI detector.

Bottom line

OpenAI’s textGrain rollout makes machine-readable text provenance more concrete, especially for EU-facing products. But the safe interpretation is narrow: a detected watermark is evidence of an OpenAI provenance signal, not proof of authorship, ownership, accuracy or responsibility. Build provenance from generation records, editorial controls and detection together rather than asking one detector to answer all of those questions.

Sources & useful resources