September 3, 2026
AI Humanizer Benchmark: September 2026 AI Humanizer Rankings
The first cycle of the benchmark is in: 11 AI humanizers, each rewriting the same 33 freshly generated texts, every output scored by 7 commercial AI detectors plus meaning preservation and readability. That is 363 rewrites and 2,541 detector verdicts, and the single most important number to come out of them is this: only 64 rewrites, 17.6% of the field's output, were judged human by all seven detectors at once. Every tool on the board got caught by something.
UndetectedGPT won the cycle at 84.9 out of 100. WriteHuman posted the field's highest raw bypass rate at 90%, and StealthGPT its most readable output at 83%. The full spread runs from 84.9 down to 38.2, and the interesting stories are in how tools earned, and lost, their positions. The complete table, with every sub-score and downloadable per-tool data, is on the leaderboard.
September 2026 leaderboard
Full interactive table| # | Humanizer | Overall | Bypass | Meaning | Read. | Pen. |
|---|---|---|---|---|---|---|
| 1– | 84.90 | 86.3 | 84.8 | 77.3 | None | |
| 2– | 84.80 | 85.3 | 87.6 | 75.7 | None | |
| 3– | 81.60 | 90.4 | 76.2 | 72.4 | −1.0 | |
| 4– | 78.30 | 79.9 | 75.7 | 76.7 | None | |
| 5– | 76.10 | 82.6 | 70.8 | 74.7 | None | |
| 6– | 73.80 | 73.2 | 80.8 | 63.8 | None | |
| 7– | 71.90 | 74.6 | 70.1 | 82.7 | −3.0 | |
| 8– | 68.50 | 70.7 | 77.3 | 74.0 | −5.0 | |
| 9– | 66.80 | 86.5 | 65.3 | 73.2 | −11.0 ×2 | |
| 10– | 39.00 | 3.7 | 59.2 | 81.4 | −4.0 ×2 | |
| 11– | 38.20 | 0.2 | 89.6 | 71.7 | −12.0 ×2 |
Overall is 0 to 100; bypass, meaning, and readability are 0 to 100 (higher is better). Penalties are points deducted for output-quality issues, already reflected in the overall score.
The tool with the most perfect scores finished ninth
CleverHumanizer produced 14 of the cycle's 64 all-seven sweeps, more than any other tool, including both tools that beat it by eighteen points. Its bypass rate, 86%, is fourth-best. It still finished ninth, because its rewrites stopped saying what the originals said: 65% meaning preservation, the second-worst in the field, with four severe-drift penalties and seven length-inflation penalties on top.
This is the exact case the scoring formula exists for. A rewrite that beats detectors by replacing the text with different text has not humanized anything; it has substituted. Bypass carries 42% of the composite precisely so that the other 58% can stop this strategy from winning, and in cycle one it did.
What separates first from third
The podium tools took the opposite path: evasion without much damage. UndetectedGPT paired 86% bypass with 85% meaning; SmartHumanizer traded one point of bypass for three points of meaning. WriteHuman posted the field's best raw evasion at 90% but gave back more text fidelity than either, at 76%. Across 33 samples, a tenth of a point decided first place, which is exactly the kind of margin that should flip between cycles. Nobody should treat rank one versus rank two as a durable fact yet; treat it as two tools in a dead heat, compared side by side here.
Last place preserved meaning best
Undetectable AI finished last at 38.2 with a bypass rate of zero: not one of its 33 rewrites was judged human by the panel's median verdict, and only 13 of its 231 individual detector verdicts cleared the midpoint at all. And yet its meaning preservation, 90%, was the best in the entire field, and its readability was fine. It also drew 19 length-inflation triggers (capped at the maximum penalty) with output running 1.44 times the input on average, and returned one text essentially unchanged.
The honest reading: on the default model every ordinary user gets, this tool is a careful rewriter that modern detectors see straight through. Its stronger models exist behind opt-in settings, and this benchmark deliberately tests defaults, because defaults are what a normal user experiences. That makes this a finding about the default configuration, not a ceiling on what the product can do, and it is the clearest illustration in the data of why we publish settings alongside scores.
The detectors disagree with each other, a lot
No detector agreed with the panel. Copyleaks was the strictest, averaging 40.4% human-likelihood across all 363 rewrites; Grammarly the most permissive at 78.9%. And the disagreements are not uniform: UndetectedGPT cleared GPTZero on 80% of its tests but Copyleaks on only 21%, while HIX Bypass showed the mirror image, 16% against GPTZero and 55% against Copyleaks.
The practical conclusion for anyone reading vendor marketing: a bypass claim proven against one detector is close to meaningless. Which detector matters as much as whether. The detector pages rank every tool against each detector separately.
Business email breaks almost everyone
Averaged across all tools, business email was humanized successfully 38.8% of the time, against 77.4% for academic essays, exactly half as often. Three tools scored at or under 1% on the category. Short, formulaic text gives a rewriter very little room to vary structure and vocabulary, which is most of what humanization is. The category anomaly belongs to StealthGPT, which posted 84% on business email, the best in the field, while ranking seventh overall. If your use case is short-form professional text, the overall ranking is not the ranking you need; the per-category views are.
Check any of this yourself
Before this cycle ran, we published a SHA-256 commitment to the secret value that selects the prompts; the value itself is now public, and re-deriving the prompt set from it is part of the standard verification. Every input, output, and detector verdict for all 363 tests is browsable on the evidence pages and downloadable from the public repository, where one command recomputes the entire leaderboard from the raw files. If any number in this post does not match what you compute, that is a finding; dispute it.
The October cycle runs on freshly generated texts next month. If you build a humanizer that should be on this board, submit it.