OpenAI has begun adding an invisible watermark to text its models generate, and has published the detection rates rather than only the claim. The method, which the company calls textGrain, biases word choices so that a detector can later look for a statistical signal in a passage.
From 5 October, API customers around the world can opt in to watermarked output for selected models. Over the coming weeks the watermark will be applied to eligible ChatGPT and Codex output for users on all plans in the EU only. OpenAI points to the EU AI Act, which it says requires generative AI providers to make generated text identifiable in a machine-readable way, and to the accompanying Code of Practice.
The numbers are the interesting part, and they are modest. At a target false positive rate of 1%, OpenAI says its detector identified watermarks in about 80% of 200-token passages, compared with about 95% of 400-token passages. Replacing 10% of words with synonyms reduced detection from about 92% to 66%.
OpenAI is also plain about what the watermark does not do. It does not measure how much a human contributed, establish ownership or legal responsibility, identify a user, or say anything about whether the text is accurate. A short passage, lightly edited, is close to undetectable.
Source: OpenAI, Our approach to EU text provenance rules.
