September 2026 prompt evidence
Each of these 33 prompts was generated fresh for this cycle, then handed to all 11 ranked humanizers, and every rewrite was scored by 7 commercial AI detectors. Open a prompt to read the input, every tool's rewrite, and each detector's verdict.
“Passed” counts the ranked tools whose rewrite of that prompt at least 6 of the 7 detectors scored as human-written (0.50 or above); “Failed” counts those flagged by at least one detector that scored them; “Hardest” is the single detector that flagged the most rewrites of that prompt.
How to read these scores
Each detector returns a human-likelihood on a common 0 to 1 scale, where 1 means it judged the text human-written and 0 means it flagged it as AI. A verdict counts as passed when that score is at least 0.50, the midpoint of the scale; hover any dot for the exact value. That threshold exists only to draw the dots: the bypass rate on the leaderboard is the mean of each test's median score across the 7 detectors, a continuous number, so a tool's pass count and its bypass rate will not be the same figure. Likewise a detector-rate row is the mean of that detector's scores, the same number the detector pages rank by, not a count of green dots.
Raw data
Every run on this page, downloadable as JSONL: one file per tool, each row carrying the source text, the rewrite, the settings used, and all 7 detector verdicts.