AI Humanizer Benchmark

AI Humanizer Leaderboard

Cycle: September 2026 · 33 samples per humanizer · how scoring works

Last tested
2026-09-03T16:11:14.775Z
Sample size
33
Methodology
v1.0.0
Humanizers
11
Cycle
Category
AI humanizer ranking for the September 2026 cycle, ordered by Overall.
RankHumanizerPenaltiesTrendLast tested
1UndetectedGPTundetectedgpt.ai84.9086.384.877.3nonetoday
2SmartHumanizersmarthumanizer.ai84.8085.387.675.7nonetoday
3WriteHumanwritehuman.ai81.6090.476.272.4-1.0today
4GPTinfgptinf.com78.3079.975.776.7nonetoday
5AI Humanizeaihumanize.io76.1082.670.874.7nonetoday
6HIX Bypassbypass.hix.ai73.8073.280.863.8nonetoday
7StealthGPTstealthgpt.ai71.9074.670.182.7-3.0today
8SuperHumanizersuperhumanizer.ai68.5070.777.374.0-5.0today
9CleverHumanizercleverhumanizer.ai66.8086.565.373.2-11.0today
10ReHumanizerehumanize.io39.003.759.281.4-4.0today
11Undetectable AIundetectable.ai38.200.289.671.7-12.0today

11 humanizers · 7 detectors · scores out of 100, higher is better. Click a column to sort; hover a penalty for the breakdown.

Where humanizers struggle

Mean bypass across all 11 tools, by writing category. Short, formulaic text is much harder to humanize than long-form prose.

Business email38.8%
Application essay60.4%
News article66.3%
Blog post67.5%
Discussion board69.9%
Marketing copy73.9%
Academic essay77.4%

Bypass vs. meaning

Beating detectors is only half the job. The upper right is where a rewrite both passes as human and still says what the original said. Hover any icon for the tool and its numbers.

0%0%25%25%50%50%75%75%100%100%bypass rate →meaning preserved →

Cycle data

Every per-test row for this cycle as JSONL. Humanizer files carry the source text, the output and the settings it ran with; detector files carry the text scored and the score returned. That is everything needed to re-derive the leaderboard: paste any published output into a detector and compare. The bundle is hash-committed before the run and publicly timestamped when it completes.

Humanizers

UndetectedGPT
33 rows · .jsonl
SmartHumanizer
33 rows · .jsonl
WriteHuman
33 rows · .jsonl
GPTinf
33 rows · .jsonl
AI Humanize
33 rows · .jsonl
HIX Bypass
33 rows · .jsonl
StealthGPT
33 rows · .jsonl
SuperHumanizer
33 rows · .jsonl
CleverHumanizer
33 rows · .jsonl
ReHumanize
33 rows · .jsonl
Undetectable AI
33 rows · .jsonl

Detectors

GPTZero
363 rows · .jsonl
Winston AI
363 rows · .jsonl
Originality.ai
363 rows · .jsonl
ZeroGPT
363 rows · .jsonl
Grammarly
363 rows · .jsonl
QuillBot
363 rows · .jsonl
Copyleaks
363 rows · .jsonl

How scores are computed

See the full methodology

Overall score formula

42%
Bypass

How often the rewritten text is classified as human by the detector panel.

32%
Meaning

How closely the rewrite preserves the original text's meaning.

16%
Readability

Whether the output is clear, fluent, natural writing.

10%
Consistency

Whether performance holds across all seven writing types or varies widely between them.

Penalties (right) subtract from the composite, and the result never goes below zero.

Possible penalties

CodePenaltyMax
  • Identical to input

    The text came back essentially unchanged; no meaningful rewriting occurred.

    -2.0
    -10.0
  • Refusal

    The tool declined to process the text, typically due to its own content filter.

    -1.0
    -10.0
  • Meaning drift

    The rewrite departed substantially from the original's meaning.

    -1.0
    -10.0
  • Length inflation

    The output ran far longer than the input, diluting the AI signal with added words rather than removing it.

    -1.0
    -10.0
  • Length deflation

    The output ran far shorter than the input; content was dropped rather than rephrased.

    -1.0
    -10.0

Maximum possible total penalty: -50.0 (off 100).