AI detector scores are easy to quote and easy to overread. A tool may report a probability, a human or AI label, or a product-specific term such as Sapling's "Fake" score. Those outputs are signals, not a record of authorship. This 2026 benchmark puts the same three English samples through six detectors and documents the visible results, test conditions, and limits of a small pilot.
Test date:
Key Findings from the 2026 Pilot
- Six external AI detectors were compared: GPTZero, Pangram, Walter Writes, ZeroGPT, Originality.ai, and Sapling.
- Five of six visible primary labels pointed toward AI-generated text for the raw AI control. ZeroGPT was the exception, showing 13.8% AI GPT while its page label said Human written.
- Five of six tools labeled the public-domain human control as human or original. ZeroGPT returned the opposite direction on this sample, which is why its numeric score and page label are both reported below.
- The AI-assisted edited sample split the tools. Three returned an AI or AI-paraphrasing result, while three returned a human or original result.
These are descriptive counts from one scan per condition, not detector accuracy rates. A proper accuracy study would need a larger, balanced dataset, repeated trials, confidence intervals, and a predefined scoring rule.
AI Detector Results: Six Tools, Three Controlled Samples
The table preserves each product's visible wording where possible. The scores are not normalized because the products expose different metrics. An AI probability, a confidence label, and Sapling's Fake metric should not be averaged as if they were the same measurement.
| Detector | Synthetic AI control about 205 words | Human control about 225 to 228 words | AI-assisted edited sample about 271 to 273 words |
|---|---|---|---|
| GPTZero Model 4.9b | AI 100% Highly confident AI generated | Human 100% Entirely human | AI 100% Possible AI Paraphrasing |
| Pangram Model 4.0 | AI 100% | Human Written 100% | AI 100% Appears to have been paraphrased or rewritten |
| Walter Writes | 96% Probability AI generated | 99% Probability Human generated | 88% Probability Human generated |
| ZeroGPT | 13.8% AI GPT Page label: Human written | 97.3% AI GPT Page label: AI/GPT Generated | 0% AI GPT Page label: Human written |
| Originality.ai Classic model, app v4.7.6 | Likely AI 100% Confident | Likely Original 100% Confident | Likely Original 99% Confident |
| Sapling | Fake: 100.0% | Fake: 0.3% | Fake: 99.9% |
Reading the table: Sapling's page uses the word "Fake" for the metric shown above, so the report keeps that label rather than silently renaming it. ZeroGPT is also reported with both its numeric AI GPT score and its page label because the two visible signals pointed in different directions in this pilot. These details are part of the observation and should not be treated as a universal product conclusion.
How We Tested
The benchmark used three fixed English prose conditions. Every detector received the same text for its condition, pasted into the public detector interface. The first visible result after the primary scan was recorded. No result was selected because it matched an expected label.
- Synthetic AI control: a 205-word sample generated for this test and left unedited. It describes how small teams use AI writing tools and why human review still matters.
- Public-domain, pre-LLM literary control: a short excerpt from Alice's Adventures in Wonderland. This is a narrow human-written control, not a representative sample of all human writing.
- AI-assisted edited sample: the synthetic AI control was processed through Ryter Pro's Humanizer and scanned as a separate condition. It is an AI-assisted transformation, not human ground truth.
- Language
- English
- Conditions
- 3 fixed samples
- External detectors
- 6 products
- Approximate sample length
- 205 to 273 words
- Primary scan date
- September 10, 2026
The sample lengths were kept above the short-input range commonly seen in quick detector demos, but they are still too small for claims about long-form publishing, academic submissions, or every genre and language. Detector versions and interfaces can change, so this page should be read as a dated benchmark snapshot.
What This Pilot Does and Does Not Show
This benchmark is useful for comparing visible behavior under one controlled setup. It is not a claim that one detector is the best, that a percentage equals a probability of misconduct, or that a single scan can establish who wrote a passage.
- It shows: detector disagreement on a fixed set of samples, including a clear split on the AI-assisted edited condition.
- It does not show: sensitivity, specificity, F1 score, false-positive rate, or product-wide performance.
- Scores can change with: text length, genre, language, model family, detector version, formatting, and post-processing.
- A responsible interpretation: a high AI score is not proof of authorship, and a human score is not proof of human authorship.
In a broader set of texts tested separately, we also observed samples that returned human or original-style results in GPTZero, Pangram, and Sapling. Those examples were kept outside this controlled table because their source texts and setup were different. Keeping them separate makes the report easier to reproduce and prevents a mixed sample from looking like a formal accuracy study.
What Published Benchmarks Show
Independent research points in the same direction: detector performance depends heavily on the data and task. A few useful reference points are summarized here without turning them into a direct ranking of commercial products.
| Reference | Design | Useful caution |
|---|---|---|
| BUST benchmark, NAACL 2024 ACL Anthology ID 2024.naacl-long.444 | About 25,000 human and LLM-generated texts, covering seven LLMs, ten tasks, and three sources; five detectors were evaluated. | Performance varied substantially across tasks and text characteristics. |
| M4GT-Bench, 2024 Wang et al., arXiv:2402.11175 | Multilingual, multi-domain, multi-generator evaluation covering binary detection, model attribution, and mixed-text boundary detection. | Strong results usually depend on training and test data sharing the same domain and generator conditions. |
| NIST AI 700-1, published 2025 2024 GenAI Text-to-Text Pilot | Human and machine summaries were evaluated with measures including AUC and Brier score across multiple systems. | Some generators were difficult for most discriminators, while some discriminators detected almost every generator in the tested setting. |
| Frontiers in Education study, 2024 | 459 unique student responses were assessed with five detectors, including GPTZero, ZeroGPT, and Originality.ai. | The study reported useful separation in its sample but also non-trivial false positives on human writing. |
Source notes: BUST, NAACL 2024, ACL Anthology ID 2024.naacl-long.444; M4GT-Bench, arXiv:2402.11175; NIST AI 700-1, 2024 NIST GenAI Pilot Study: Text-to-Text Evaluation Overview and Results; Frontiers in Education, 2024, DOI 10.3389/feduc.2024.1374889.
How to Use Detector Results Responsibly
- Keep provenance with the text. Record whether the sample was human-written, generated, translated, edited, or transformed, and preserve the exact version used for testing.
- Report the detector's own language. Keep the raw score and visible label together. Do not turn different vendor metrics into one blended percentage.
- Use multiple forms of evidence. For a high-stakes review, combine detector output with drafts, revision history, source checking, the writer's explanation, and ordinary editorial judgment.
- Repeat after material changes. A new model, a new detector version, or a major rewrite creates a new test condition. If a workflow is used to humanize AI text, report that transformation instead of presenting the result as an untouched human sample.
For content teams, the practical value of a detector is often comparative rather than judicial. It can help document how a workflow behaves across different drafts, but it should not be used to make an unsupported claim about a writer's identity or intent.
How to Cite This Benchmark
Ryter Pro, "AI Detector Benchmark 2026: A 6-Tool Test," tested September 10, 2026. Controlled pilot of three English samples across GPTZero, Pangram, Walter Writes, ZeroGPT, Originality.ai, and Sapling; n=1 per condition.
This wording gives a reader the date, scope, products, and sample size in one place. If the page is updated, keep the stable URL and add a short changelog so older references remain understandable.
Frequently Asked Questions
What is an AI detector benchmark?
It is a documented comparison of one or more AI detection tools against defined text samples. A useful benchmark states the sample source, language, length, test date, detector version when available, and the rule used to interpret each output.
Can AI detectors prove who wrote a text?
No. A detector result is a model output about text characteristics. It can support a review, but it cannot by itself prove authorship, intent, or academic misconduct.
Why do AI detector scores disagree?
Products use different models, training data, thresholds, labels, and input rules. Results also change with language, genre, length, editing, translation, and the generator that produced the text.
How often should this benchmark be updated?
Review it when a major detector model or interface changes, and at least annually for a page positioned as a 2026 reference. Keep the URL stable, date every result, and add new rows or conditions without silently replacing the old record.
Summary
This 2026 pilot shows why provenance and context matter more than a single percentage. Six detectors agreed more often on the two control samples than on the AI-assisted edited sample, while one product displayed a numeric score and page label that pointed in different directions. Use the table as a dated reference, keep the raw wording visible, and update it when models or detector versions change. Teams that document an AI writing or editing workflow can use Ryter Pro as one part of that record, while keeping detector results in their proper role: evidence to review, not a verdict.
