# Perplexity, Burstiness, and the Myth of the Machine Fingerprint

> Every AI detector claims to see something in your writing. Here is what they are actually measuring — and why it was never a fingerprint to begin with.

- Author: [Mara Okafor](https://www.detectionguide.com/authors/mara-okafor/), Editor-in-chief
- Published: 2026-10-01
- Section: Detectors
- URL: https://www.detectionguide.com/articles/perplexity-burstiness-explained/

## Key takeaways

- Perplexity measures how predictable text is to a language model. Low perplexity looks machine-made to a detector.
- Burstiness measures how much sentence length, structure and predictability vary across a document. Human writing tends to vary more.
- Most commercial detectors also use trained classifiers, zero-shot statistical tests such as DetectGPT and Binoculars, or watermarks.
- None of these is a fingerprint: predictable human writing gets flagged, light paraphrasing defeats detection, and short text is unreliable.

Ask a detector vendor how their product works and you will usually hear two words: perplexity and burstiness. They sound like lab equipment. They are closer to weather forecasting — useful, probabilistic, and wrong often enough that you should never bet a student's semester on them.

This is a plain-language tour of what detectors measure, where the idea came from, and why the "machine fingerprint" so many people imagine does not exist.

## Perplexity: how surprised is the model?

A language model is, at heart, a machine for guessing the next word. Feed it "The cat sat on the" and it will assign a high probability to "mat" and a very low one to "spreadsheet."

**Perplexity** turns that guessing into a score. Run a passage through a model and ask, word by word, how likely each choice was. If the text keeps picking the words the model expected, perplexity is low. If it keeps zigging where the model expected a zag, perplexity is high.

The intuition behind detection follows naturally: text generated by a model tends to be made of high-probability choices, because that is how the model produced it. So low perplexity looks machine-made. Human writers, the theory goes, are weirder.

## Burstiness: does the rhythm change?

Perplexity is measured across a whole passage, but humans are not consistently weird. We write a long, winding sentence full of qualifications. Then a short one. Then something mid-length that wanders off.

**Burstiness** tries to capture that unevenness: how much predictability, sentence length, and structure vary across a document. GPTZero, one of the first consumer detectors, built its early pitch around exactly these two signals.

You can see the crudest version of the idea below. This toy only measures how much sentence length varies — nothing about word probabilities — but it shows why "smooth" text gets flagged.

## The other families of detectors

Perplexity and burstiness are the famous ones, but most commercial tools today blend several approaches:

- **Trained classifiers.** Take a large pile of text labeled "human" and "AI," then train a model to tell them apart. OpenAI released a detector for GPT-2 output built this way back in 2019. These work well on the kind of text they were trained on and degrade on everything else.
- **Zero-shot statistical tests.** Methods like DetectGPT (2023) look at how a model's probability for a passage changes when the passage is slightly perturbed; Binoculars (2024) compares how two different models see the same text. No training data needed — but they still depend on access to a model similar to the one that wrote the text.
- **Watermarks.** If the company generating the text deliberately bakes in a statistical signal, detection becomes much more reliable. That only works when the generator cooperates, which is its own story.

> A detector doesn’t find evidence of a machine. It finds evidence of predictability — and plenty of humans are predictable.
>
> — The core problem

## Why it was never a fingerprint

A fingerprint is unique and stable. Detector signals are neither.

**Predictable humans exist.** Legal boilerplate, lab reports, five-paragraph essays, and anything written to a tight template have low perplexity by design. So does writing by people working in a second language, who often lean on safer, more common phrasing. Researchers at Stanford showed in 2023 that several popular detectors flagged essays by non-native English speakers as AI-generated at alarming rates.

**Famous text is "predictable" too.** Anything that appeared thousands of times in a model's training data — scripture, constitutions, famous speeches — looks extremely likely to that model. That is why screenshots of detectors flagging the US Constitution went viral.

**Light editing changes everything.** Researchers have repeatedly shown that paraphrasing AI output, even with another AI, can drop detection rates sharply. That is the entire business model of the humanizer industry.

**Short text is noise.** A paragraph does not contain enough words for the statistics to settle. Many vendors quietly recommend minimum lengths for exactly this reason.

> **What this means for you**
>
> A detector score is a probability estimate from a model with known blind spots. Treat it like a smoke alarm, not a
> courtroom: a reason to look closer, never proof on its own.

## So are detectors useless?

Not quite. On long, unedited, purely machine-generated text, good detectors do catch a lot. The trouble is that the real world rarely serves up that clean case. It serves up a student who brainstormed with a chatbot, wrote their own draft, and ran it through a grammar checker — and then gets a number with two decimal places and no error bars.

That gap between how detectors are built and how they are used is what this publication exists to cover.

## Common questions

### What is perplexity in AI detection?

Perplexity is a score for how predictable a passage is to a language model, measured word by word. Text made of the words a model expected has low perplexity, which detectors treat as a sign of machine generation.

### What is burstiness in AI detection?

Burstiness captures how much a document varies in sentence length, structure and predictability. Human writers tend to mix long and short sentences, while model output often holds a steadier rhythm.

### Can AI detectors reliably tell whether a human wrote something?

Not reliably. Predictable human writing such as templates, legal text and second-language writing can be flagged, light editing can drop detection rates sharply, and short passages don’t contain enough words for the statistics to settle. Treat a score as a reason to look closer, never as proof.
