# How We Test Detectors and Humanizers

> Corpora, thresholds, blind readers and the numbers we refuse to publish without context. The DetectionGuide testing protocol, in full.

- Author: [Jules Abernathy](https://www.detectionguide.com/authors/jules-abernathy/), Lab lead
- Published: 2026-09-18
- Updated: 2026-09-28
- Section: Lab
- URL: https://www.detectionguide.com/articles/how-we-test/

## Key takeaways

- We report every detector’s false positive rate alongside its detection rate, at a threshold fixed before testing.
- Our corpora cover pre-2022 human writing, contemporary and second-language human writing, unedited AI output, and hybrid human–AI text.
- Humanizers are judged on evasion across several independent detectors and on quality rated by blind human readers.
- Every result is date-stamped, and vendors never see results before publication.

Every week, someone sends us a screenshot: a detector saying 0% AI, or 100%, or some oddly specific number in between. Screenshots are not tests. This page explains what *is* a test at DetectionGuide, so you can judge our results — and anyone else's — on their merits.

## Principles

**Measure errors, not just hits.** A detector that flags everything catches all AI text. It also convicts every student. We always report false positive rates alongside detection rates.

**Fix the threshold before testing.** Many tools let you choose how aggressive to be. We record the setting and keep it fixed across a test; we don't pick whichever setting makes the result look most interesting.

**Date-stamp everything.** Detectors and humanizers update constantly, often without notice. A result is a snapshot of a specific product on a specific day.

**No vendor previews.** Companies don't see results before publication, and we don't accept free access in exchange for coverage.

## Our corpora

We maintain several sets of text, each designed to probe a different failure mode:

| Set | What it contains | What it tests |
| --- | --- | --- |
| Pre-2022 human | Writing published before modern chatbots were widely available | Baseline false positives |
| Contemporary human | Recent writing contributed with consent, with drafting history | Real-world false positives |
| Second-language | Writing by non-native English speakers, with consent | Known bias |
| Pure generated | Output from several current models, unedited | Baseline detection |
| Hybrid | Human drafts with AI edits, and AI drafts with human edits | The messy middle |

Contributors are told how their writing will be used and can withdraw it at any time.

## Testing detectors

For each detector we report, at a fixed threshold:

- **True positive rate:** of the AI-generated texts, how many were flagged.
- **False positive rate:** of the human-written texts, how many were wrongly flagged — broken out by corpus.
- **Length sensitivity:** how both rates change for short texts versus long ones.

Where a tool provides a continuous score, we also report how well it separates the two classes across all thresholds. But the headline is always the false positive rate at the setting people actually use.

## Testing humanizers

A humanizer that defeats detectors by producing nonsense isn't a success. So we test two things.

**Evasion.** We run generated text through the humanizer and measure how detection rates change across several independent detectors — not the humanizer's own built-in checker.

**Quality.** Blind human readers compare original and humanized versions without knowing which is which, rating clarity, accuracy and naturalness. We also check whether the meaning changed, especially hedges, numbers and technical terms.

## What we won't publish

- Single screenshots presented as findings.
- Results from samples too small to support the claim.
- Rankings without error bars or sample sizes.
- Any test of a tool on a student's real submission without their consent.

## Corrections

If you think we got something wrong, tell us. If we did, we'll fix it, date the fix, and say what changed at the top of the article.

## Common questions

### How should an AI detector be tested fairly?

Fix the threshold before testing, use both AI-generated and human-written text, and report the false positive rate alongside the detection rate, broken out by type of writing and by text length. Date-stamp the results, because tools change without notice.

### How do you measure whether an AI humanizer works?

We measure evasion as the change in detection rates across several independent detectors, not the humanizer’s own checker, and quality by having blind human readers rate clarity, accuracy and naturalness and checking whether the meaning changed.
