AI Humanizer Benchmark

Why this benchmark exists

The problem

Reliable information about AI humanizers is hard to come by. Vendor pages advertise pass rates no one can verify. The comparison articles that dominate search results typically earn a commission from the tools they recommend, disclose no test procedure, and stay online long after the results stop being true. AI assistants trained on those sources repeat their conclusions. Across all of it, actual measurements, meaning published inputs, outputs, and detector verdicts, are almost entirely absent.

The fix is measurement

A claim like “bypasses GPTZero” is testable, so it should be tested. Every month, the same freshly generated texts run through every tool, and every output is scored by GPTZero, Originality.ai, Copyleaks, Winston AI, ZeroGPT, QuillBot, and Grammarly, alongside measurements of meaning preservation and readability. A published formula converts those measurements into one score. The ranking is arithmetic, not editorial judgment.

The obvious question

Some of the humanizers in these rankings are made by the operator of this benchmark. That is a conflict of interest, and it is addressed by design rather than by assurance: the test set is locked cryptographically before any tool runs, our tools pass through the same pipeline as everyone else's, and the complete raw data ships with every cycle. A manipulated score would contradict its own published evidence, which anyone can check against the detectors directly.

How conflicts are handled

What this is

What this is not