AI Humanizer Benchmark

Leaderboard / Compare / StealthGPT vs WriteHuman

Head to head · September 2026 cycle

StealthGPT vs WriteHuman

WriteHuman finished 9.7 points ahead of StealthGPT in the September 2026 cycle, 81.60 to 71.90, ranking #3 against #7 in a field of 11. WriteHuman scored higher on 6 of 7 detectors and 6 of 7 writing categories. StealthGPT came out ahead on Readability (82.7 vs 72.4), GPTZero (78.2 vs 56.5), Business email (81.3 vs 69.5). The bypass-rate gap, 74.6 to 90.4, sits inside both tools' 95% confidence intervals, so treat bypass as a draw.

71.90/100
Overall score
#7/11
Rank
+9.7
WriteHuman leads
Detectors
16
Categories
16
Components
14
WriteHumanLeads
writehuman.ai
81.60/100
Overall score
#3/11
Rank
Last tested
2026-09-03
Prompts
33
Methodology
v1.0.0

Score components

What the overall score is made of: bypass rate weighs 42%, meaning preservation 32%, readability 16%, consistency across categories 10%. Output-quality penalties come off the total. How scoring works

StealthGPTWriteHuman
Bypass rateCIs overlapWriteHuman by 15.8
74.6
90.4
Meaning preservationWriteHuman by 6.1
70.1
76.2
ReadabilityStealthGPT by 10.3
82.7
72.4
ConsistencyWriteHuman by 8.0
78.5
86.4
Penaltiesfewer is betterWriteHuman by 2.0
−3.0
−1.0

CIs overlap marks a gap that sits inside both tools' 95% confidence intervals for bypass rate; it may not survive another cycle.

Detector by detector

Share of each tool's outputs that a detector classified as human-written. WriteHuman took 6 of 7.

StealthGPTWriteHuman
GPTZeroStealthGPT by 21.8
78.2
56.5
Winston AIWriteHuman by 20.0
62.1
82.1
Originality.aiWriteHuman by 11.0
48.2
59.2
ZeroGPTWriteHuman by 25.5
66.5
92.0
GrammarlyWriteHuman by 16.6
77.5
94.2
QuillBotWriteHuman by 26.3
66.0
92.2
CopyleaksWriteHuman by 6.1
57.6
63.6

Writing categories

Category score per writing context: bypass credit only for real rewrites, blended with the category's own meaning preservation and readability. Small categories swing; the prompt count is shown on each row.

Academic essayApplication essayBlog postBusiness emailMarketing copyDiscussion boardNews article
StealthGPTWriteHumanOuter ring = 100
StealthGPTWriteHuman
Academic essay6 promptsWriteHuman by 6.2
80.0
86.2
Application essay6 promptsWriteHuman by 18.4
66.5
84.9
Blog post6 promptsWriteHuman by 4.0
76.6
80.6
Business email3 promptsStealthGPT by 11.8
81.3
69.5
Marketing copy6 promptsWriteHuman by 1.2
85.0
86.2
Discussion board3 promptsWriteHuman by 25.9
57.2
83.1
News article3 promptsWriteHuman by 10.8
68.8
79.6

Where each one wins

Every measure above that one tool won outright, largest margin first within each group.

StealthGPT

3 wins
  • Readability Score82.7 vs 72.4
  • GPTZero Detector78.2 vs 56.5
  • Business email Category81.3 vs 69.5

WriteHuman

16 wins
  • Bypass rate Score90.4 vs 74.6
  • Consistency Score86.4 vs 78.5
  • Meaning preservation Score76.2 vs 70.1
  • Penalties Score−1.0 vs −3.0
  • QuillBot Detector92.2 vs 66.0
  • ZeroGPT Detector92.0 vs 66.5
  • Winston AI Detector82.1 vs 62.1
  • Grammarly Detector94.2 vs 77.5
  • Originality.ai Detector59.2 vs 48.2
  • Copyleaks Detector63.6 vs 57.6
  • Discussion board Category83.1 vs 57.2
  • Application essay Category84.9 vs 66.5
  • News article Category79.6 vs 68.8
  • Academic essay Category86.2 vs 80.0
  • Blog post Category80.6 vs 76.6
  • Marketing copy Category86.2 vs 85.0

Questions people ask

Is StealthGPT better than WriteHuman?

WriteHuman finished 9.7 points ahead of StealthGPT in the September 2026 cycle, 81.60 to 71.90, ranking #3 against #7 in a field of 11. WriteHuman scored higher on 6 of 7 detectors and 6 of 7 writing categories. Both tools ran the same 33 prompts, scored by the same 7 detectors, under methodology v1.0.0; every input, output, and verdict is published.

Which bypasses AI detectors better, StealthGPT or WriteHuman?

WriteHuman posted the higher bypass rate, 90.4 vs 74.6, and scored higher on 6 of 7 detectors. StealthGPT led on GPTZero; WriteHuman led on Winston AI, Originality.ai, ZeroGPT, Grammarly, QuillBot and Copyleaks. The overall bypass gap is inside both 95% confidence intervals, so it is not a reliable difference.

Do StealthGPT and WriteHuman keep the original meaning?

WriteHuman preserved meaning better, 76.2 vs 70.1 on our 0–100 similarity scale. On readability, StealthGPT rated higher, 82.7 vs 72.4. Penalties this cycle: StealthGPT −3.0, WriteHuman −1.0.

More matchups

Add a third tool to this comparison·Every matchup

Go deeper