AI Humanizer Benchmark

September 3, 2026

Introducing AI Humanizer Benchmark

Shop for an AI humanizer and you will meet the same number everywhere: a bypass rate somewhere north of 99%, set in large type, with no test behind it. The comparison articles ranking above the vendors are mostly affiliate pages that earn a commission on their own first pick. The reviews below those are frequently seeded. A buyer who works through this stack carefully ends up knowing one thing: somebody wants a sale. Which tool actually gets AI text past AI detectors, and at what cost to the text, is unknowable from any of it.

It should not be. A claim like "bypasses GPTZero" is a measurable proposition. So we measure it, monthly, for every major humanizer at once, and publish everything the measurement produces.

What a cycle is

Once a month, a fresh set of 33 texts is generated across seven kinds of writing: academic essays, application essays, blog posts, business emails, marketing copy, discussion posts, and news articles. Every humanizer on the roster rewrites the identical 33 texts on its own default settings and base plan, because defaults are what an ordinary user actually gets. Every rewrite then goes to seven commercial AI detectors: GPTZero, Originality.ai, Copyleaks, Winston AI, ZeroGPT, QuillBot, and Grammarly. Alongside the detector verdicts we measure whether the rewrite still says what the original said, and whether it is still readable.

A published formula turns those measurements into one score: bypass 42%, meaning 32%, readability 16%, consistency across categories 10%, minus penalties for quality failures like padding the text or handing it back unchanged. No editorial adjustment happens anywhere. If the formula says a tool is fourth, the page says fourth.

The part that makes it checkable

Publishing a leaderboard is easy; publishing one that can be audited is the actual work. Every cycle ships its complete raw record: all inputs, all 363 outputs, every one of the 2,541 detector verdicts, and the exact scoring code that produced the ranking, in a public repository under open licenses. One command recomputes the entire leaderboard from those files. The evidence pages show the same record in browsable form, prompt by prompt, and any single verdict can be checked with no tooling at all: paste a published output into the detector and compare.

The test set itself is protected from us. Prompts are selected by a secret random value whose SHA-256 hash is published before the cycle runs; the value is revealed when the cycle closes. Anyone can confirm the prompts follow from the committed value, which means nobody, including us, can pick prompts after seeing how tools perform on them, and no vendor can pre-compute next month's test to train against it.

The conflict, stated plainly

Some tools in these rankings are made by this benchmark's operator. That is a conflict of interest, and it gets an engineering answer rather than a promise: the same prompts, settings policy, detectors, scoring code, and penalties apply to every tool; the test set is cryptographically fixed before any tool runs; and the complete raw data is published, so favoritism would have to survive public recomputation of every number. No vendor pays for placement, no vendor supplies us access, and there is not an affiliate link on the site. The longer version lives on the fairness page.

What the first cycle already shows

The September 2026 cycle is published, and it is a strong argument for measuring instead of trusting: only 64 of 363 rewrites passed all seven detectors, the tool with the most perfect sweeps finished ninth because it mangled meaning, and the sharpest tool on paper scored 0% bypass on its default settings. The full analysis is in the September recap, and the numbers behind every sentence of it are on the leaderboard.

If you build a humanizer

Get on the board: submit your tool and it joins the next cycle under the identical rules. If you believe a published score is wrong, dispute it with a test id; every row ships with its input and output, so most disagreements resolve by looking at the data. A new cycle lands every month, and this blog will carry a recap each time, along with any methodology changes, which are versioned and never retroactive.

← All posts