What should “undetectable AI” mean when detectors disagree and writing quality still matters? We tested two humanizer candidates on an untouched holdout, measured detector scores and factual fidelity separately, and published the misses alongside the wins.
- 12
- frozen source topics
- 32
- holdout rewrites
- 2
- public detectors
- 0%
- strict holdout pass
Research question
Can a detector-free humanizer reconstruct complete AI-style drafts in a less canonical form while preserving facts, readability, and genre fit better than the former generic rewrite prompt?
We separated three outcomes because a low detector score is not useful when a rewrite loses facts or becomes awkward:
- Detector result: below 50% AI on every detector used for that comparison.
- Factual fidelity: at least 95% fact recall, with no altered material facts or unsupported material claims.
- Writing quality: at least 4 out of 5 for meaning preservation, readability, naturalness, and genre fit.
How the test was designed
- We froze four development topics and eight holdout topics before final scoring.
- Every source draft had to score at least 90% AI on ZeroGPT. The 12 accepted baselines averaged 99.35%.
- The final holdout compared two candidates across eight untouched topics with two generations per topic: 16 outputs per candidate and 32 in total.
- ZeroGPT and AI Detector were measured in a browser session on July 11, 2026.
- A separate evaluation model scored the frozen fact requirements, unsupported claims, meaning, readability, naturalness, and genre fit.
- Historical controls were rescored to expose detector drift rather than assuming the services were stable.
Frozen-holdout results
| Candidate | Quality pass | Average fact recall | Average ZeroGPT AI | Average AI Detector AI | Strict pass |
|---|---|---|---|---|---|
| Lower-cost finalist | 56.25% | 98.23% | 41.08% | 88.81% | 0% |
| Selected Stealth candidate | 93.75% | 100% | 46.39% | 83.94% | 0% |
How to read this: the candidates averaged below 50% on ZeroGPT but remained high on AI Detector. An average is not a pass rate, and the 0% strict result is the most important detector finding.
What the holdout changed
Development results made the reconstruction strategy look broadly effective. The untouched holdout challenged that conclusion. Structural rewriting reduced one detector's score, but the second detector remained high. The selected candidate still earned its production role because it preserved the required facts and writing quality far more reliably—not because it defeated every detector.
The strongest practical technique was rebuilding from an internal fact-and-qualifier ledger instead of replacing words sentence by sentence. A final audit for omitted facts, invented claims, and strengthened qualifiers mattered more than generic instructions to “sound human.” Multi-stage repair and deliberately rough grammar did not generalize well.
Detector drift showed up in the controls
Six saved historical outputs were rescored. ZeroGPT reproduced all six earlier scores exactly. AI Detector moved in both directions: one control rose from 28% to 73%, while another fell from 19% to 3%. The text did not change; the service result did.
This is why detector scores should be recorded with a date and treated as measurements of a particular service at a particular time. They are not permanent facts about a document's authorship.
What “undetectable AI” should mean in practice
The phrase is often used as if one rewrite could guarantee a human label everywhere. Our results do not support that promise. A more defensible workflow is to use an undetectable AI detector and humanizer workspace to identify generic passages, rebuild structure, preserve the factual payload, and then review the result as a person.
A detector can still be useful as a diagnostic signal. It should not replace source checking, revision history, a writer's explanation, or human judgment. Our guide to how AI detectors work explains that distinction in more detail.
Limitations
- The holdout contained eight topics and two generations per candidate per topic.
- Only two browser-accessible detectors were used.
- A separate evaluation model, rather than a human panel, scored the factual-quality gate.
- Generation is stochastic, and both models and detector services can change after the experiment date.
- Famous textbook topics were especially difficult because common facts tend to appear in conventional language and order.
- These results describe this experiment; they are not a claim that every customer input will behave the same way.
Download the aggregate data
The public dataset contains the test design, thresholds, baseline summary, holdout aggregates, historical controls, conclusion, and limitations. It excludes customer text, personal information, credentials, and proprietary prompt text.
Download the 2026 benchmark results as JSON
Research context
Our test was informed by published work showing why evaluation must include multiple domains, detector drift, paraphrase robustness, false-positive tradeoffs, and text quality:
- RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors (ACL 2024)
- A Practical Examination of AI-Generated Text Detectors for Large Language Models (NAACL Findings 2025)
- Can AI-Generated Text be Reliably Detected?
- Paraphrasing Evades Detectors of AI-Generated Text, but Retrieval Is an Effective Defense
Use the result as a signal, then review the writing
Check AI-likelihood, reconstruct weak passages, and compare the final draft for facts, meaning, tone, and reader fit.
Open the AI humanizer