# Your Essay Isn’t a Robot. Neither Is the Constitution.

> False positives aren’t a rounding error in AI detection. They are the predictable result of how detectors work — and the people they land on are not random.

- Author: [Mara Okafor](https://www.detectionguide.com/authors/mara-okafor/), Editor-in-chief
- Published: 2026-09-29
- Section: Detectors
- URL: https://www.detectionguide.com/articles/false-positives-constitution/

## Key takeaways

- AI detectors flag text that a language model finds predictable, so heavily quoted human writing such as the US Constitution can be scored as AI-generated.
- A 2023 Stanford study in the journal Patterns found detectors misclassified more than half of TOEFL essays by non-native English speakers as AI-generated, on average.
- Base rates matter: at a 1% false positive rate, 200,000 honest essays would produce about 2,000 false flags.
- A detector flag should start a conversation, not end one.

In mid-2023, a screenshot made the rounds: a popular AI detector, fed the text of the United States Constitution, confidently declared it AI-generated. It was funny, briefly. Then people realized the joke had a punchline, and it was aimed at students.

The Constitution was not flagged because of a bug. It was flagged because the detector was working exactly as designed.

## Why famous text looks "fake"

Most detectors lean on some version of a simple question: *how predictable is this text to a language model?* Text the model finds very predictable gets pushed toward the "AI" end of the scale.

But a document that appears thousands of times across the web — quoted, analyzed, reprinted, taught — is about as predictable as text gets. The model has effectively memorized it. To the detector, memorized human writing and freshly generated machine writing look alike: both are things the model would have said itself.

The Constitution is an extreme case. The everyday version is subtler and much more consequential.

## Who gets flagged

In 2023, a team of researchers at Stanford published a study in the journal *Patterns* testing several widely used GPT detectors on two sets of essays: one written by US eighth-graders, the other written by non-native English speakers for the TOEFL exam.

The detectors did reasonably well on the eighth-grade essays. On the TOEFL essays, they misclassified more than half as AI-generated on average — and nearly every TOEFL essay in the sample was flagged by at least one of the detectors.

The reason is the same one that sank the Constitution. People writing in a second language often favor common words and conventional structures. That is good, careful writing practice. It is also low perplexity.

Other groups pay the same tax:

- **Students taught to formula.** The five-paragraph essay is a template. Templates are predictable.
- **Technical and legal writers.** Precision means repetition. Repetition means predictability.
- **Neurodivergent writers** who prefer consistent structure and explicit transitions.
- **Anyone who used a grammar tool**, which nudges prose toward the most common, "correct" phrasing.

## The base-rate problem

Even a detector with a low false positive rate produces a lot of false accusations at scale. That isn't an opinion; it's arithmetic.

| Scenario | Honest essays | False positive rate | Wrongly flagged |
| --- | --- | --- | --- |
| One course section | 120 | 1% | ~1 |
| A large first-year class | 2,000 | 1% | ~20 |
| A university, one term | 200,000 | 1% | ~2,000 |

*Illustrative figures. Real-world false positive rates depend heavily on the detector, the threshold and the kind of writing.*

And 1% is optimistic for many of the populations above. Vendors tend to publish rates measured on their own test sets, at their own thresholds, on long documents. Classroom reality — short responses, mixed drafting, second-language writers — is messier.

## What a flag should mean

None of this means detection is pointless. It means a flag is the *start* of a conversation, not the end of one. Most detector vendors say as much in their own documentation, usually in smaller type than the score itself.

The problem is that a percentage with two decimal places looks like evidence. It feels like a lab result. And people with heavy workloads and little training on how these systems fail are being handed that number and asked to make decisions with it.

The Constitution survived its AI accusation. Most students don't get a viral screenshot to vouch for them.

## Common questions

### Why did an AI detector say the US Constitution was written by AI?

Detectors score how predictable text is to a language model. A document quoted and reprinted thousands of times is extremely predictable to a model that has effectively memorized it, so it looks the same as freshly generated text.

### Are AI detectors biased against non-native English speakers?

Research suggests so. A 2023 Stanford study published in Patterns found that detectors misclassified more than half of TOEFL essays by non-native English speakers as AI-generated on average, while doing reasonably well on essays by US eighth-graders.

### How many students can a low false positive rate affect?

At a 1% false positive rate, about 1 in a section of 120 honest essays, 20 in a class of 2,000, and 2,000 across a university term of 200,000 would be wrongly flagged. Real rates vary by detector, threshold and type of writing, and 1% is optimistic for many groups.
