detectionguide
Subscribe

Share

Detectors · Explainer

Perplexity, Burstiness, and the Myth of the Machine Fingerprint

Every AI detector claims to see something in your writing. Here is what they are actually measuring — and why it was never a fingerprint to begin with.

4 min read
Author
Mara OkaforOctober 1, 20262:20 PM UTC
Section
DetectorsSection
Length
4 minutesReading time
HUMAN · SENTENCE LENGTHMODEL · SENTENCE LENGTHDG/861
Human sentences swing between short and long. Model output tends to hold a steady rhythm. Detectors bet on the difference.

Ask a detector vendor how their product works and you will usually hear two words: perplexity and burstiness. They sound like lab equipment. They are closer to weather forecasting — useful, probabilistic, and wrong often enough that you should never bet a student’s semester on them.

This is a plain-language tour of what detectors measure, where the idea came from, and why the “machine fingerprint” so many people imagine does not exist.

Perplexity: how surprised is the model?

A language model is, at heart, a machine for guessing the next word. Feed it “The cat sat on the” and it will assign a high probability to “mat” and a very low one to “spreadsheet.”

Perplexity turns that guessing into a score. Run a passage through a model and ask, word by word, how likely each choice was. If the text keeps picking the words the model expected, perplexity is low. If it keeps zigging where the model expected a zag, perplexity is high.

The intuition behind detection follows naturally: text generated by a model tends to be made of high-probability choices, because that is how the model produced it. So low perplexity looks machine-made. Human writers, the theory goes, are weirder.

Burstiness: does the rhythm change?

Perplexity is measured across a whole passage, but humans are not consistently weird. We write a long, winding sentence full of qualifications. Then a short one. Then something mid-length that wanders off.

Burstiness tries to capture that unevenness: how much predictability, sentence length, and structure vary across a document. GPTZero, one of the first consumer detectors, built its early pitch around exactly these two signals.

You can see the crudest version of the idea below. This toy only measures how much sentence length varies — nothing about word probabilities — but it shows why “smooth” text gets flagged.

Interactive · Burstiness meter

Sentences—
Avg. words—
Variation—
Reads as—
This only measures sentence-length variation (coefficient of variation). Real detectors use language-model probabilities, and even they get it wrong. Nothing you type leaves your browser.

The other families of detectors

Perplexity and burstiness are the famous ones, but most commercial tools today blend several approaches:

  • Trained classifiers. Take a large pile of text labeled “human” and “AI,” then train a model to tell them apart. OpenAI released a detector for GPT-2 output built this way back in 2019. These work well on the kind of text they were trained on and degrade on everything else.
  • Zero-shot statistical tests. Methods like DetectGPT (2023) look at how a model’s probability for a passage changes when the passage is slightly perturbed; Binoculars (2024) compares how two different models see the same text. No training data needed — but they still depend on access to a model similar to the one that wrote the text.
  • Watermarks. If the company generating the text deliberately bakes in a statistical signal, detection becomes much more reliable. That only works when the generator cooperates, which is its own story.

A detector doesn’t find evidence of a machine. It finds evidence of predictability — and plenty of humans are predictable.

— The core problem

Why it was never a fingerprint

A fingerprint is unique and stable. Detector signals are neither.

Predictable humans exist. Legal boilerplate, lab reports, five-paragraph essays, and anything written to a tight template have low perplexity by design. So does writing by people working in a second language, who often lean on safer, more common phrasing. Researchers at Stanford showed in 2023 that several popular detectors flagged essays by non-native English speakers as AI-generated at alarming rates.

Famous text is “predictable” too. Anything that appeared thousands of times in a model’s training data — scripture, constitutions, famous speeches — looks extremely likely to that model. That is why screenshots of detectors flagging the US Constitution went viral.

Light editing changes everything. Researchers have repeatedly shown that paraphrasing AI output, even with another AI, can drop detection rates sharply. That is the entire business model of the humanizer industry.

Short text is noise. A paragraph does not contain enough words for the statistics to settle. Many vendors quietly recommend minimum lengths for exactly this reason.

So are detectors useless?

Not quite. On long, unedited, purely machine-generated text, good detectors do catch a lot. The trouble is that the real world rarely serves up that clean case. It serves up a student who brainstormed with a chatbot, wrote their own draft, and ran it through a grammar checker — and then gets a number with two decimal places and no error bars.

That gap between how detectors are built and how they are used is what this publication exists to cover.

Common questions

What is perplexity in AI detection?

Perplexity is a score for how predictable a passage is to a language model, measured word by word. Text made of the words a model expected has low perplexity, which detectors treat as a sign of machine generation.

What is burstiness in AI detection?

Burstiness captures how much a document varies in sentence length, structure and predictability. Human writers tend to mix long and short sentences, while model output often holds a steadier rhythm.

Can AI detectors reliably tell whether a human wrote something?

Not reliably. Predictable human writing such as templates, legal text and second-language writing can be flagged, light editing can drop detection rates sharply, and short passages don’t contain enough words for the statistics to settle. Treat a score as a reason to look closer, never as proof.

  • #perplexity
  • #burstiness
  • #how-it-works

Written by

Mara Okafor

Mara has spent a decade covering the strange places where software meets judgment. She reads model cards for fun and has never once trusted a confidence score.

More from Mara

Keep reading

All stories
detectionguide
Subscribe

The arms race between the machines that write and the machines that judge.

↑↓ to move↵ to open/ to search anywhere