detectionguide
Subscribe

Share

Lab · Methodology

How We Test Detectors and Humanizers

Corpora, thresholds, blind readers and the numbers we refuse to publish without context. The DetectionGuide testing protocol, in full.

3 min read
Author
Jules AbernathySeptember 18, 20261:00 PM UTC
Updated
UpdatedSeptember 28, 2026
Section
LabSection
Length
3 minutesReading time
39%HUMANAICONFIDENCE · UNCALIBRATEDDG/773
A score without a threshold, a sample size and a date is not a result. It is a vibe.

Every week, someone sends us a screenshot: a detector saying 0% AI, or 100%, or some oddly specific number in between. Screenshots are not tests. This page explains what is a test at DetectionGuide, so you can judge our results — and anyone else’s — on their merits.

Principles

Measure errors, not just hits. A detector that flags everything catches all AI text. It also convicts every student. We always report false positive rates alongside detection rates.

Fix the threshold before testing. Many tools let you choose how aggressive to be. We record the setting and keep it fixed across a test; we don’t pick whichever setting makes the result look most interesting.

Date-stamp everything. Detectors and humanizers update constantly, often without notice. A result is a snapshot of a specific product on a specific day.

No vendor previews. Companies don’t see results before publication, and we don’t accept free access in exchange for coverage.

Our corpora

We maintain several sets of text, each designed to probe a different failure mode:

Set What it contains What it tests
Pre-2022 human Writing published before modern chatbots were widely available Baseline false positives
Contemporary human Recent writing contributed with consent, with drafting history Real-world false positives
Second-language Writing by non-native English speakers, with consent Known bias
Pure generated Output from several current models, unedited Baseline detection
Hybrid Human drafts with AI edits, and AI drafts with human edits The messy middle

Contributors are told how their writing will be used and can withdraw it at any time.

Testing detectors

For each detector we report, at a fixed threshold:

  • True positive rate: of the AI-generated texts, how many were flagged.
  • False positive rate: of the human-written texts, how many were wrongly flagged — broken out by corpus.
  • Length sensitivity: how both rates change for short texts versus long ones.

Where a tool provides a continuous score, we also report how well it separates the two classes across all thresholds. But the headline is always the false positive rate at the setting people actually use.

Testing humanizers

A humanizer that defeats detectors by producing nonsense isn’t a success. So we test two things.

Evasion. We run generated text through the humanizer and measure how detection rates change across several independent detectors — not the humanizer’s own built-in checker.

Quality. Blind human readers compare original and humanized versions without knowing which is which, rating clarity, accuracy and naturalness. We also check whether the meaning changed, especially hedges, numbers and technical terms.

What we won’t publish

  • Single screenshots presented as findings.
  • Results from samples too small to support the claim.
  • Rankings without error bars or sample sizes.
  • Any test of a tool on a student’s real submission without their consent.

Corrections

If you think we got something wrong, tell us. If we did, we’ll fix it, date the fix, and say what changed at the top of the article.

Common questions

How should an AI detector be tested fairly?

Fix the threshold before testing, use both AI-generated and human-written text, and report the false positive rate alongside the detection rate, broken out by type of writing and by text length. Date-stamp the results, because tools change without notice.

How do you measure whether an AI humanizer works?

We measure evasion as the change in detection rates across several independent detectors, not the humanizer’s own checker, and quality by having blind human readers rate clarity, accuracy and naturalness and checking whether the meaning changed.

  • #methodology
  • #testing
  • #transparency

Written by

Jules Abernathy

Jules runs the DetectionGuide test bench: corpora, protocols, spreadsheets, and the occasional argument about what a false positive rate actually means.

More from Jules

Keep reading

All stories
detectionguide
Subscribe

The arms race between the machines that write and the machines that judge.

↑↓ to move↵ to open/ to search anywhere