AI
AI Detector Accuracy in 2026: We Retested the Big Three
Our July 2026 retest of Originality.ai, GPTZero and JustDone. All three lost accuracy, newest models pass as human, and one now flags genuine human writing.
We retested the three most-used AI content detectors in July 2026 against current-generation models. All three had lost accuracy since our earlier testing, and they now fail in different and equally serious directions. Originality.ai remains the most accurate at 88 to 90% on clearly AI-written text, down from 96%, but it has become over-sensitive and now flags some fully human writing as AI. GPTZero fell to 79 to 85% and frequently reads output from the newest models as human. JustDone came last at 59 to 61%, barely better than a coin toss on short text.
This page is the summary of that testing. The full method and per-tool findings are in the three linked reviews.
Results at a glance
| Detector | Earlier testing | 2026 retest | How it now fails |
|---|---|---|---|
| Originality.ai | ~96% | 88 to 90% | Over-sensitive: flags some 100% human writing as AI |
| GPTZero | up to 90% | 79 to 85% | Under-sensitive: misses newest models (Fable 5, GPT-5.6 Sol) |
| JustDone AI | ~70% | 59 to 61% | Weakest overall, unreliable under 100 words |
The three findings that matter
1. The newest models defeat detection
Text produced by current-generation models, including Fable 5 and GPT-5.6 Sol, was frequently read as human. Detectors work by measuring statistical regularity, chiefly perplexity, meaning how predictable the text is, and burstiness, meaning how much sentence length and structure vary. Older models wrote smoothly and evenly, which left a clear fingerprint. Newer models vary their rhythm the way people do, so the fingerprint is faint or absent. This is the structural reason accuracy is falling, and it will keep falling as models improve.
2. Turning up sensitivity punishes real people
Originality.ai appears to have been tuned more aggressively to keep catching newer models. It is still the most accurate tool we tested, but the cost is over-sensitivity: in our 2026 testing it labelled genuinely human writing as AI, with no pattern the writer could have avoided. That is the detector trap in one sentence. Catch more AI and you catch more innocent humans. There is no setting that avoids both.
3. Editing defeats every tool
Accuracy is highest on raw AI output and falls as a human edits toward their own voice. When we rewrote AI text with the newest models and deliberately seeded human-style mistakes, the best tool dropped to about 85%. The people most likely to be caught by a detector are the ones who made the least effort to hide anything, which is the opposite of what an enforcement tool should do.
What this means for schools and universities
This is the part we think matters most. A tool that is wrong somewhere between 10% and 40% of the time, depending on the tool and the content, is not evidence. It cannot support an academic integrity charge on its own, and the failure modes fall hardest on writing that is already at risk of being misread: short answers, highly technical or formulaic prose, and writing by non-native English speakers. Meanwhile a student using a current model can pass through the same system untouched. Any policy that treats a detector score as proof will produce false accusations against honest students while missing the cases it was built to catch.
Our recommendation is unchanged and now stronger: use detectors as a prompt for a human conversation, never as a verdict. Ask to see drafts and version history. Weigh process evidence over any score.
How we tested
We ran the same categories of text through each detector: output from current-generation models, older AI output, lightly edited AI, heavily edited and paraphrased AI, and genuine human writing across casual, academic, and technical styles. We tested short samples under 100 words, mid-length pieces, and long pieces over 1500 words, and we recorded correct detections, false positives (human flagged as AI), and false negatives (AI missed). Results are directional rather than laboratory-grade, and detector behaviour changes as vendors retrain, which is precisely why we repeat this.
Full per-tool method and findings: Originality.ai, GPTZero, JustDone AI.
Where detection goes next
After-the-fact detection is losing, and it is unlikely to recover. Every model generation narrows the statistical gap the tools depend on. The industry is already shifting toward provenance, proving where content came from rather than guessing afterwards, through content credentials and watermarking standards, disclosure norms where AI assistance is declared rather than hunted, and process-based proof such as version history and draft trails. Detectors will probably survive as one weak signal inside that mix. The era of treating a detector score as evidence is ending.
Citing this study
Journalists, educators, and researchers are welcome to cite these findings with attribution to MPG ONE and a link to this page. If you need the underlying breakdown or want to discuss the method, contact us. We plan to repeat this testing as models and detectors change, so the figures here carry a test date and will be updated rather than quietly edited.
FAQ
What is the most accurate AI detector in 2026?
Originality.ai was the most accurate in our 2026 testing at 88 to 90% on clearly AI-written content, ahead of GPTZero at 79 to 85% and JustDone at 59 to 61%. It has also become over-sensitive, so it sometimes flags human writing as AI.
Can AI detectors detect the newest AI models?
Often not. Output from current-generation models such as Fable 5 and GPT-5.6 Sol was frequently read as human, especially by GPTZero. Newer models produce the natural variation that detectors use to identify human writing.
Do AI detectors falsely accuse students?
They can. In our 2026 testing, Originality.ai flagged some fully human writing as AI, and false positives cluster on short, technical, formulaic, and non-native English writing. No detector score should be treated as proof of misconduct.
Are AI detectors getting better or worse?
Worse. Every tool we tested lost accuracy compared with our earlier testing, because each new model generation writes with more human-like variation and erodes the signal detectors rely on.
Helpful next steps
Turn the idea into something useful
When this topic becomes part of a real website, workflow, campaign, or AI system, these MPG ONE pages are the natural places to continue.
Build generative AI workflows for content, documents, knowledge, and brand-safe operations.
AI development servicesBuild private AI assistants, workflow automation, knowledge systems, and reporting layers.
AI agent development servicesDesign tool-using AI agents with memory, retrieval, permissions, and human review.
Custom AI solutionsCreate a custom AI system around your data, tools, permissions, and operating model.
Marketing and SEO servicesConnect technical SEO, content strategy, campaigns, analytics, and conversion-focused pages.