AI Humanizer Benchmark

Detectors / GPTZero

Best for GPTZero

Humanizers ranked by their measured bypass rate against GPTZero. This cycle it was the 4th-strictest of the 7 detectors in the panel: humanizers averaged 53.9% bypass against it.

The detector most widely used in education, where much AI-checking of written work actually happens.

Reports a 0-1 AI probability; inverted to human = 1 - ai.

Last tested
2026-09-03T16:11:14.775Z
Sample size
33
Methodology
v1.0.0
AI humanizer ranking for the September 2026 cycle, ordered by GPTZero bypass.
RankHumanizerPenaltiesLast tested
1CleverHumanizercleverhumanizer.ai90.066.80-11.0today
2UndetectedGPTundetectedgpt.ai80.084.90nonetoday
3StealthGPTstealthgpt.ai78.271.90-3.0today
4SmartHumanizersmarthumanizer.ai71.084.80nonetoday
5GPTinfgptinf.com69.878.30nonetoday
6SuperHumanizersuperhumanizer.ai64.068.50-5.0today
7WriteHumanwritehuman.ai56.581.60-1.0today
8AI Humanizeaihumanize.io55.476.10nonetoday
9HIX Bypassbypass.hix.ai15.773.80nonetoday
10Undetectable AIundetectable.ai6.138.20-12.0today
11ReHumanizerehumanize.io5.739.00-4.0today

11 humanizers · 7 detectors · scores out of 100, higher is better. Click a column to sort; hover a penalty for the breakdown.

Raw test data

Every verdict GPTZero returned in the September 2026 cycle as JSONL: the text scored and the score, one row per test. Paste any row's text back into the detector to check the recorded verdict.

GPTZero
363 rows · .jsonl

How scores are computed

See the full methodology

Overall score formula

42%
Bypass

How often the rewritten text is classified as human by the detector panel.

32%
Meaning

How closely the rewrite preserves the original text's meaning.

16%
Readability

Whether the output is clear, fluent, natural writing.

10%
Consistency

Whether performance holds across all seven writing types or varies widely between them.

Penalties (right) subtract from the composite, and the result never goes below zero.

Possible penalties

CodePenaltyMax
  • Identical to input

    The text came back essentially unchanged; no meaningful rewriting occurred.

    -2.0
    -10.0
  • Refusal

    The tool declined to process the text, typically due to its own content filter.

    -1.0
    -10.0
  • Meaning drift

    The rewrite departed substantially from the original's meaning.

    -1.0
    -10.0
  • Length inflation

    The output ran far longer than the input, diluting the AI signal with added words rather than removing it.

    -1.0
    -10.0
  • Length deflation

    The output ran far shorter than the input; content was dropped rather than rephrased.

    -1.0
    -10.0

Maximum possible total penalty: -50.0 (off 100).