detectionguide
Subscribe

Share

Detectors · Analysis

Watermarks Were Supposed to Settle This

Hide a statistical signature in every sentence a model writes, and detection becomes math instead of guesswork. So why hasn’t it ended the argument?

3 min read
Author
Jules AbernathySeptember 22, 202610:00 AM UTC
Section
DetectorsSection
Length
3 minutesReading time
GREEN-LIST TOKENS · Z = 5.7DG/156
A watermark nudges a model toward a secret “green list” of tokens. Invisible to readers, obvious to anyone holding the key.

Every argument about AI detection eventually arrives at the same hopeful idea: what if the model just told us?

Not with a disclaimer at the bottom of the text, which anyone can delete, but with a signal woven into the words themselves. That’s the promise of text watermarking. It’s real, it works in the lab, and it still hasn’t settled anything.

How a text watermark works

The best-known approach comes from a 2023 paper by researchers at the University of Maryland. The idea is elegant.

Every time a model picks the next word, it chooses from a ranked list of candidates. A watermarking scheme uses a secret key to split the vocabulary into a “green list” and a “red list” at each step, then gently nudges the model toward green words. A reader notices nothing: there are usually several perfectly good words to pick from.

But anyone holding the key can count green words. Ordinary human text lands near the 50/50 split you’d expect by chance. Watermarked text leans green far more often than chance allows, and with enough words the statistics become overwhelming.

Instead of asking “does this feel like a machine,” you ask a question with a precise answer: “how unlikely is this many green words, if a human wrote it?”

Watermarks in the wild

This is no longer just a paper. Google DeepMind has deployed a watermarking system called SynthID across its own products, and in 2024 it published the approach and released the text version as open-source tooling so other developers could use it.

OpenAI, meanwhile, was reported in 2024 to have built a text watermarking method internally and debated releasing it. The company said publicly that it was weighing the tradeoffs, including how easily the watermark could be removed and the risk of stigmatizing legitimate uses of AI — such as people using it to help write in a second language.

Why it hasn’t ended the argument

It only works if the generator cooperates. A watermark is a feature the model provider chooses to add. Open-weight models that anyone can run locally won’t carry one unless the person running them wants it to.

There’s no universal detector. Each scheme has its own key. Checking for one company’s watermark tells you nothing about text from another company’s model.

Rewriting erodes it. Because the signal lives in word choices, changing the words weakens it. Light edits leave plenty of signal; heavy paraphrasing, translation into another language and back, or a second model rewriting the text can wash much of it out. Humanizers, in other words, are a watermark-removal service whether they advertise it or not.

Short text is still hard. A watermark is a statistical claim. A two-sentence answer doesn’t contain enough words to make it confidently.

Watermarking turns detection from a guess into a measurement — but only for the text that was watermarked to begin with.

What it would take

Watermarks are most useful as one layer in a broader system of provenance: signals that travel with content and say where it came from. That system would need widespread adoption by model providers, shared standards for verification, and honest communication about what a missing watermark means. (Answer: almost nothing.)

Until then, a watermark is a strong yes in a world full of unknowns. It can tell you a particular tool wrote something. It can’t tell you that a person did.

Common questions

How does AI text watermarking work?

At each step, a secret key splits the vocabulary into a “green list” and a “red list,” and the model is gently nudged toward green words. Human text lands near a 50/50 split by chance; watermarked text leans green far more often, which anyone holding the key can measure.

Can AI text watermarks be removed?

They can be weakened. Light edits leave plenty of signal, but heavy paraphrasing, translation into another language and back, or a second model rewriting the text can wash much of it out.

Does a missing watermark mean a human wrote the text?

No. Open-weight models run locally won’t carry a watermark unless the person running them adds one, and checking for one company’s watermark tells you nothing about text from another company’s model.

  • #watermarking
  • #synthid
  • #research

Written by

Jules Abernathy

Jules runs the DetectionGuide test bench: corpora, protocols, spreadsheets, and the occasional argument about what a false positive rate actually means.

More from Jules

Keep reading

All stories
detectionguide
Subscribe

The arms race between the machines that write and the machines that judge.

↑↓ to move↵ to open/ to search anywhere