AI Humanizer Leaderboard
Cycle: September 2026 · 33 samples per humanizer · how scoring works
- Last tested
- 2026-09-03T16:11:14.775Z
- Sample size
- 33
- Methodology
- v1.0.0
- Humanizers
- 11
| Rank | Humanizer | Penalties | Trend | Last tested | ||||
|---|---|---|---|---|---|---|---|---|
| 1– | 84.90 | 86.3 | 84.8 | 77.3 | none | today | ||
| 2– | 84.80 | 85.3 | 87.6 | 75.7 | none | today | ||
| 3– | 81.60 | 90.4 | 76.2 | 72.4 | -1.0 | today | ||
| 4– | 78.30 | 79.9 | 75.7 | 76.7 | none | today | ||
| 5– | 76.10 | 82.6 | 70.8 | 74.7 | none | today | ||
| 6– | 73.80 | 73.2 | 80.8 | 63.8 | none | today | ||
| 7– | 71.90 | 74.6 | 70.1 | 82.7 | -3.0 | today | ||
| 8– | 68.50 | 70.7 | 77.3 | 74.0 | -5.0 | today | ||
| 9– | 66.80 | 86.5 | 65.3 | 73.2 | -11.0 | today | ||
| 10– | 39.00 | 3.7 | 59.2 | 81.4 | -4.0 | today | ||
| 11– | 38.20 | 0.2 | 89.6 | 71.7 | -12.0 | today |
11 humanizers · 7 detectors · scores out of 100, higher is better. Click a column to sort; hover a penalty for the breakdown.
Where humanizers struggle
Mean bypass across all 11 tools, by writing category. Short, formulaic text is much harder to humanize than long-form prose.
Bypass vs. meaning
Beating detectors is only half the job. The upper right is where a rewrite both passes as human and still says what the original said. Hover any icon for the tool and its numbers.
Cycle data
Every per-test row for this cycle as JSONL. Humanizer files carry the source text, the output and the settings it ran with; detector files carry the text scored and the score returned. That is everything needed to re-derive the leaderboard: paste any published output into a detector and compare. The bundle is hash-committed before the run and publicly timestamped when it completes.
Humanizers
Detectors
How scores are computed
See the full methodologyOverall score formula
How often the rewritten text is classified as human by the detector panel.
How closely the rewrite preserves the original text's meaning.
Whether the output is clear, fluent, natural writing.
Whether performance holds across all seven writing types or varies widely between them.
Penalties (right) subtract from the composite, and the result never goes below zero.
Possible penalties
- Identical to input
The text came back essentially unchanged; no meaningful rewriting occurred.
-2.0-10.0 - Refusal
The tool declined to process the text, typically due to its own content filter.
-1.0-10.0 - Meaning drift
The rewrite departed substantially from the original's meaning.
-1.0-10.0 - Length inflation
The output ran far longer than the input, diluting the AI signal with added words rather than removing it.
-1.0-10.0 - Length deflation
The output ran far shorter than the input; content was dropped rather than rephrased.
-1.0-10.0
Maximum possible total penalty: -50.0 (off 100).