What should “undetectable AI” mean when detectors disagree and writing quality still matters? We tested two humanizer candidates on an untouched holdout, measured detector scores and factual fidelity separately, and published the misses alongside the wins.
- 12
- frozen source topics
- 32
- holdout rewrites
- 2
- public detectors
- 0%
- strict holdout pass
Research question
Can a detector-free humanizer reconstruct complete AI-style drafts in a less canonical form while preserving facts, readability, and genre fit better than the former generic rewrite prompt?
We separated three outcomes because a low detector score is not useful when a rewrite loses facts or becomes awkward:
- Detector result: below 50% AI on every detector used for that comparison.
- Factual fidelity: at least 95% fact recall, with no altered material facts or unsupported material claims.
- Writing quality: at least 4 out of 5 for meaning preservation, readability, naturalness, and genre fit.
How the test was designed
- We froze four development topics and eight holdout topics before final scoring.
- Every source draft had to score at least 90% AI on ZeroGPT. The 12 accepted baselines averaged 99.35%.
- The final holdout compared two candidates across eight untouched topics with two generations per topic: 16 outputs per candidate and 32 in total.
- ZeroGPT and AI Detector were measured in a browser session on July 11, 2026.
- A separate evaluation model scored the frozen fact requirements, unsupported claims, meaning, readability, naturalness, and genre fit.
- Historical controls were rescored to expose detector drift rather than assuming the services were stable.
Frozen-holdout results
| Candidate | Quality pass | Average fact recall | Average ZeroGPT AI | Average AI Detector AI | Strict pass |
|---|---|---|---|---|---|
| Lower-cost finalist | 56.25% | 98.23% | 41.08% | 88.81% | 0% |
| Selected Stealth candidate | 93.75% | 100% | 46.39% | 83.94% | 0% |
How to read this: the candidates averaged below 50% on ZeroGPT but remained high on AI Detector. An average is not a pass rate, and the 0% strict result is the most important detector finding.
What the holdout changed
Development results made the reconstruction strategy look broadly effective. The untouched holdout challenged that conclusion. Structural rewriting reduced one detector's score, but the second detector remained high. The selected candidate still earned its production role because it preserved the required facts and writing quality far more reliably—not because it defeated every detector.
The strongest practical technique was rebuilding from an internal fact-and-qualifier ledger instead of replacing words sentence by sentence. A final audit for omitted facts, invented claims, and strengthened qualifiers mattered more than generic instructions to “sound human.” Multi-stage repair and deliberately rough grammar did not generalize well.
Detector drift showed up in the controls
Six saved historical outputs were rescored. ZeroGPT reproduced all six earlier scores exactly. AI Detector moved in both directions: one control rose from 28% to 73%, while another fell from 19% to 3%. The text did not change; the service result did.
This is why detector scores should be recorded with a date and treated as measurements of a particular service at a particular time. They are not permanent facts about a document's authorship.
What “undetectable AI” should mean in practice
The phrase is often used as if one rewrite could guarantee a human label everywhere. Our results do not support that promise. A more defensible workflow is to use an undetectable AI detector and humanizer workspace to identify generic passages, rebuild structure, preserve the factual payload, and then review the result as a person.
A detector can still be useful as a diagnostic signal. It should not replace source checking, revision history, a writer's explanation, or human judgment. Our guide to how AI detectors work explains that distinction in more detail.
Limitations
- The holdout contained eight topics and two generations per candidate per topic.
- Only two browser-accessible detectors were used.
- A separate evaluation model, rather than a human panel, scored the factual-quality gate.
- Generation is stochastic, and both models and detector services can change after the experiment date.
- Famous textbook topics were especially difficult because common facts tend to appear in conventional language and order.
- These results describe this experiment; they are not a claim that every customer input will behave the same way.
Full-length essay examples: September 2026
These are actual, unedited Basic and Stealth outputs from September 18, 2026, using the Essay option and the same original 442-word synthetic source. They are two selected examples from a 20-output editorial evaluation across five writing types, not a detector benchmark or an estimate of typical results. No customer writing was used.
The review checked factual details, uncertainty, completeness, and voice. Across the full evaluation, some outputs still dropped the fictional label or shortened the source more than requested. These selected examples retain the proposed study's status and its limitations; the complete passage is available below for comparison.
Read the original essay (442 words)
A school considering a change to its writing lessons should distinguish time spent away from a draft from time spent avoiding it. In a fictional six-week pilot, two classes would write their first drafts in class. One class would revise on the same afternoon, while the other would return to the work after two days. Both classes would use the same rubric and receive feedback from the same teacher. The proposal does not include test results, and it should not be treated as evidence that either schedule is already better.
The case for a delay rests on a familiar difficulty in revision. A writer knows what a paragraph was supposed to mean and may read that intended meaning into sentences that do not actually express it. Returning later can make an unexplained link or missing qualification easier to notice. A student might discover that a paragraph announces a reason but provides no example, or that a conclusion makes a stronger claim than the evidence supports. In those situations, distance from the draft may help the writer see a problem that quick proofreading would miss.
There are limits to this argument. Waiting does not provide knowledge that the writer never had, nor does it replace reading the source material carefully. If a student's account of an event is wrong, a later revision may simply make the incorrect account more fluent. Revision therefore needs questions about the argument as well as questions about wording. What does this claim depend on? Which part of the source supports it? Is an exception important enough to change the conclusion? These questions require active checking, whether revision happens immediately or later.
The proposed comparison would also be difficult to interpret. Absences could affect how much time students receive in class, and unequal help outside school could influence the second draft. A stronger piece of writing would not automatically show that the delay caused the improvement. The teacher would need to record these limitations and compare the changes made, not just the final appearance of the papers. The pilot could be useful for deciding how to organize lessons without supporting a broad claim about every student or every kind of assignment.
A delayed revision session is worth considering as part of a writing routine, but it should not crowd out planning or source checking. Students still need to decide what they want to argue, gather suitable evidence, and identify what that evidence cannot establish. The practical question is not whether waiting is a cure for weak writing. It is whether the timetable gives students a genuine opportunity to return to an argument and examine what is on the page.
Read the Basic rewrite (414 words)
A school weighing a change to its writing lessons may find that the real issue is not revision itself, but the kind of distance revision gives. The source imagines a fictional six-week pilot in which two classes write first drafts in class. One class revises the same afternoon; the other returns after two days. Both use the same rubric and receive feedback from the same teacher. No test results are included, and the proposal should not be treated as evidence that either schedule is already better.
That delay rests on a familiar problem. Writers often know what a paragraph was supposed to mean and read that intended meaning into sentences that do not actually express it. With time away from the draft, an unexplained link or missing qualification can be easier to spot. A student may see that a paragraph announces a reason but gives no example, or that a conclusion makes a stronger claim than the evidence supports. In those cases, distance may reveal a problem that quick proofreading would miss.
The argument has limits. Waiting does not create knowledge the writer never had, and it does not replace careful reading of the source material. If a student’s account of an event is wrong, later revision may only make the incorrect account more fluent. Revision therefore needs questions about the argument, not only questions about wording: What does this claim depend on? Which part of the source supports it? Is an exception important enough to change the conclusion? Those checks still have to be active, whether revision happens at once or later.
That makes the comparison hard to read. Absences could change how much class time students receive, and unequal help outside school could shape the second draft. A stronger paper would not automatically show that the delay caused the improvement. The teacher would need to record those limits and compare the changes made, not just the final appearance of the papers. Even then, the pilot could only help with how to organize lessons; it would not support a broad claim about every student or every kind of assignment.
A delayed revision session can belong in a writing routine, but it should not crowd out planning or source checking. Students still need to decide what they want to argue, gather suitable evidence, and identify what that evidence cannot establish. The practical question is whether the timetable gives them a genuine chance to return to an argument and examine what is on the page.
Read the Stealth rewrite (430 words)
The most useful measure in a proposed change to writing lessons may not be the polish of the final paper. It may be the changes students actually make after returning to a draft. In a fictional six-week pilot, two classes would write first drafts in class under the same conditions: the same rubric and feedback from the same teacher. One class would revise that afternoon; the other would return after two days. The proposal includes no test results, so it cannot show that either schedule is already better.
A stronger second draft would also require careful interpretation. Absences could alter the amount of class time available to students, while unequal help outside school could affect the later version. Even if one group produced better-looking papers, that would not automatically mean that the delay caused the improvement. The teacher would need to record those limitations and compare revisions themselves rather than judging only the final appearance of the work. Such a pilot could help organize lessons, but it could not support a broad conclusion about every student or every type of assignment.
The appeal of waiting lies in a specific difficulty. Writers often know what a paragraph was meant to say and can read that intended meaning into sentences that do not convey it. Time away from a draft can make an unexplained connection or an omitted qualification easier to see. A student may notice that a paragraph offers a reason without providing an example, or that a conclusion claims more than the available evidence supports. Distance may reveal problems that quick proofreading misses.
That benefit depends on distinguishing time away from a draft from time spent avoiding it. Waiting does not create knowledge a writer never possessed, and it does not substitute for careful reading of source material. A student whose account of an event is wrong may return later and merely make the incorrect account sound smoother. Revision must therefore address the argument as well as the wording: What does the claim depend on? Which part of the source supports it? Does an exception matter enough to alter the conclusion? Those questions call for active checking whether revision occurs immediately or after a delay.
A delayed session may have a place in a writing routine, provided it does not displace planning or source checking. Students still need to decide what they want to argue, gather suitable evidence, and identify what that evidence cannot establish. The practical issue is whether the timetable gives them a genuine chance to return to an argument and examine what is actually on the page.
Download the source, outputs, model identifiers, and source hashes. The July detector experiment above remains a separate, frozen test.
Download the aggregate data
The public dataset contains the test design, thresholds, baseline summary, holdout aggregates, historical controls, conclusion, and limitations. It excludes customer text, personal information, credentials, and proprietary prompt text.
Download the 2026 benchmark results as JSON
Research context
Our test was informed by published work showing why evaluation must include multiple domains, detector drift, paraphrase robustness, false-positive tradeoffs, and text quality:
- RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors (ACL 2024)
- A Practical Examination of AI-Generated Text Detectors for Large Language Models (NAACL Findings 2025)
- Can AI-Generated Text be Reliably Detected?
- Paraphrasing Evades Detectors of AI-Generated Text, but Retrieval Is an Effective Defense
Use the result as a signal, then review the writing
Check AI-likelihood, reconstruct weak passages, and compare the final draft for facts, meaning, tone, and reader fit.
Open the AI humanizer