AI Humanizer Benchmark

September 2026 prompt evidence

Each of these 33 prompts was generated fresh for this cycle, then handed to all 11 ranked humanizers, and every rewrite was scored by 7 commercial AI detectors. Open a prompt to read the input, every tool's rewrite, and each detector's verdict.

33
Prompts
11
Tools ranked
363
Humanizer outputs
2541
Detector verdicts
PromptCategoryInput wordsPassedFailedHardest
01Argumentative essayAcademic essay31948Originality.ai
02Lit reviewAcademic essay29369Originality.ai
03Cover letterApplication essay286011GPTZero
04Personal statementApplication essay29966GPTZero
05How-to blogBlog post281310Originality.ai
06Listicle blogBlog post298610Originality.ai
07Business emailBusiness email233111Winston AI
08Landing copyMarketing copy261710Originality.ai
09Product descriptionMarketing copy188210ZeroGPT
10Discussion postDiscussion board26769Originality.ai
11News articleNews article299310Copyleaks
12Argumentative essayAcademic essay293311Originality.ai
13Lit reviewAcademic essay30167Originality.ai
14Cover letterApplication essay281011Winston AI
15Personal statementApplication essay310410GPTZero
16How-to blogBlog post300511Originality.ai
17Listicle blogBlog post27466Originality.ai
18Business emailBusiness email227111Winston AI
19Landing copyMarketing copy24849Originality.ai
20Product descriptionMarketing copy19477GPTZero
21Discussion postDiscussion board26439Copyleaks
22News articleNews article301210GPTZero
23Argumentative essayAcademic essay31358Originality.ai
24Lit reviewAcademic essay29267Originality.ai
25Cover letterApplication essay282011Copyleaks
26Personal statementApplication essay30877Copyleaks
27How-to blogBlog post311210Originality.ai
28Listicle blogBlog post29169Copyleaks
29Business emailBusiness email223011GPTZero
30Landing copyMarketing copy27069Originality.ai
31Product descriptionMarketing copy18394GPTZero
32Discussion postDiscussion board26438Copyleaks
33News articleNews article28279GPTZero

“Passed” counts the ranked tools whose rewrite of that prompt at least 6 of the 7 detectors scored as human-written (0.50 or above); “Failed” counts those flagged by at least one detector that scored them; “Hardest” is the single detector that flagged the most rewrites of that prompt.

How to read these scores

Each detector returns a human-likelihood on a common 0 to 1 scale, where 1 means it judged the text human-written and 0 means it flagged it as AI. A verdict counts as passed when that score is at least 0.50, the midpoint of the scale; hover any dot for the exact value. That threshold exists only to draw the dots: the bypass rate on the leaderboard is the mean of each test's median score across the 7 detectors, a continuous number, so a tool's pass count and its bypass rate will not be the same figure. Likewise a detector-rate row is the mean of that detector's scores, the same number the detector pages rank by, not a count of green dots.

Raw data

Every run on this page, downloadable as JSONL: one file per tool, each row carrying the source text, the rewrite, the settings used, and all 7 detector verdicts.

UndetectedGPT
33 rows · .jsonl
SmartHumanizer
33 rows · .jsonl
WriteHuman
33 rows · .jsonl
GPTinf
33 rows · .jsonl
AI Humanize
33 rows · .jsonl
HIX Bypass
33 rows · .jsonl
StealthGPT
33 rows · .jsonl
SuperHumanizer
33 rows · .jsonl
CleverHumanizer
33 rows · .jsonl
ReHumanize
33 rows · .jsonl
Undetectable AI
33 rows · .jsonl