Every AI detector sells you a percentage. Winston AI says 99.98 percent accuracy. GPTZero says 99 percent. Originality.ai says 97.8 percent for its multilingual model. Turnitin, more usefully, publishes a false positive rate of 0.013.
Those numbers all sound like the problem is solved. Then you multiply one of them by a class size, and it stops sounding that way.
That multiplication is the whole point of this post, because nearly every article about AI detectors either repeats the vendor percentage or repeats a claim about bias, and almost none of them do the arithmetic that turns a percentage into a number of people.
The number that matters is not accuracy
Accuracy is the wrong headline figure, and it is the one every vendor leads with.
An AI detector can be wrong in two directions. It can miss AI text, which is a false negative, or it can flag human writing as AI, which is a false positive. Those two errors are not remotely equal in cost. A false negative means a piece of generated text slips through and nobody notices. A false positive means a person is told they cheated.
One of those is a missed catch. The other is an accusation with a name attached.
So when a detector advertises a single accuracy number, it is blending the error you care about with the error you mostly do not. The figure to ask for is the false positive rate, and most of these companies do not put it on the marketing page.
What the vendors actually claim
Here is what each one publishes on its own site, read today.
| Tool | Claimed accuracy | Publishes a false positive rate |
|---|---|---|
| Winston AI | 99.98 percent | No figure, says low false positives |
| GPTZero | 99 percent, 96.5 percent on mixed documents | States it trains ESL false positives down to 1 percent |
| Originality.ai | 97.8 percent multilingual | Yes, in a published accuracy study |
| Turnitin | Cites independent research | Yes, 0.013 and 0.014 |
| ZeroGPT | No figure published | No |
| Sapling | No figure published | Says plainly that false positives will occur |
Look at the top two rows. Winston claims an error rate of 0.02 percent. GPTZero claims 1 percent. Those are the same task, and they differ by a factor of fifty.
They cannot both be right. There is no shared test set behind these numbers, so each company is measuring its own thing and calling it accuracy. Treat the percentages as marketing until somebody shows you the test.
Now do the multiplication
Take the most credible published figure, which is Turnitin's, because it is specific and it comes with a method. Call it 1.3 percent of human documents wrongly flagged.
A single essay at 1.3 percent feels like nothing. You would take that bet.
A class of 300 essays produces about four falsely flagged students.
A department running 5,000 submissions in a year produces about 65.
A university processing 50,000 submissions a year produces roughly 650 students whose honest work was flagged as machine written.
Set out as a table, because this is the part worth screenshotting.
| Documents scanned | Falsely flagged at 1.3 percent | Roughly |
|---|---|---|
| 1 essay | 0.013 | Nothing to worry about |
| 30, one class | 0.4 | Less than one a term |
| 300, a course | 3.9 | About four students |
| 5,000, a department | 65 | More than one a week |
| 50,000, a university | 650 | Two every working day |
None of those students did anything wrong. The rate did not change. Only the denominator did, and the denominator is the part nobody puts on the pricing page.
The bottom row is the one to sit with. At a mid-sized university, on the most credible published figure in this whole category, roughly two honest students are flagged every working day of the year. That is not a scandal or a bug report. It is the advertised performance of the product working exactly as specified.
This is the same shape as every other number I have taken apart on this site. Runway's free tier is a real offer until you convert credits into seconds. Gamma's free plan is generous until you notice the meter runs per slide. A 1.3 percent error rate is negligible until you multiply it by how many times you run the test.
A score is a probability, not a verdict
Every one of these tools returns a number, and almost everyone reads that number as an answer. It is not one.
Originality.ai makes the point better than I can, and it works against them commercially, so it is worth quoting the substance: a result of 60 percent likely original and 40 percent likely AI is not automatically a false positive. The tool is expressing uncertainty, not identifying 40 percent of the document as machine written.
That distinction collapses the moment the score reaches somebody without the training to read it. A number between 0 and 100 next to a student's name looks like a measurement. It gets treated like a breathalyser reading, when it is closer to a weather forecast.
There is a second trap in how the scores get used. If a teacher runs a document through three different detectors and acts on the highest score, they have not triangulated anything. They have selected for whichever tool was most wrong in the direction they were already leaning. Three tools with independent error rates produce more false positives between them than any one of them alone, not fewer.
If you take one operational rule from this post, make it this: one tool, one threshold, agreed in advance, and the score never travels without the word count beside it.
The bias claim, and why I could not confirm it
I started this expecting to write that AI detectors are biased against non-native English speakers. It is the most repeated criticism of these tools, it spread from early research on older detectors, and it is stated as settled fact in a great many articles.
Turnitin's own published evaluation does not support it. They tested nearly 2,000 English Language Learner writing samples and report a 0.014 false positive rate for ELL writers against 0.013 for native English writers. That is not a meaningful gap.
I want to be careful about how much weight that carries. This is a vendor publishing research about its own product, on its own website, alongside a citation to independent work in which it performs well. That is not the same as an independent replication, and you should discount it accordingly.
But it is specific, it names a sample size, and it is checkable. That puts it ahead of a bare 99.98 percent with no methodology attached. So my honest position is that the bias claim is unproven on current tools rather than established, and that anyone stating it flatly in 2026 is repeating something they have not checked.
GPTZero, for what it is worth, implicitly concedes the concern exists. Its own site says it repeatedly retrains to push the ESL false positive rate down to 1 percent, which is both an admission that the problem was real and a rate seven times higher than Turnitin's.
The 300 word threshold changes everything
This is the most practically useful thing I found, and it is buried.
Turnitin's published false positive figures apply to documents that meet a 300 word minimum. That is not a footnote. It means the reliability numbers simply do not describe shorter work.
Sapling goes further and says the more essay-like a piece of writing is, the more likely a false positive becomes.
Put those together. Short answers, discussion board posts, single paragraph reflections and brief written responses are the least reliable case for these tools, and they are exactly the work that gets scanned in bulk because there is so much of it.
If you are a teacher, this gives you a defensible policy in one line: do not run detection on anything under 300 words, because the vendor's own numbers do not cover it.
The most honest page any of them publish
Sapling's detector page includes this: false positives and false negatives will occur.
No percentage, no qualifier, no claim of being the most accurate detector on the market. It then explains that formal, essay-like writing is more likely to trigger a false positive, which is an admission directly against its own commercial interest.
That single sentence is more useful than every 99 percent claim in this post combined, because it tells you how to hold the output. Not as a verdict. As a probability that will sometimes be wrong in the expensive direction.
What free AI detectors actually give you
Most of the well-known detectors run a free tier that caps how much text you can check at once, with the full document handling behind a paid plan. GPTZero, ZeroGPT, Sapling, Winston AI and Originality.ai all still operate and all responded when I checked them today.
Two things worth knowing before you rely on any of them.
Copyleaks, Quillbot and Scribbr all block automated access, so I could not read their current terms directly and I am not going to quote you numbers I could not verify. And Writer.com's dedicated AI content detector page now redirects to an anchor on its homepage rather than a page of its own, which is usually what happens shortly before a feature stops being a priority.
If you want the wider picture of which free AI tools still hold up, my best free AI tools roundup covers the same checking method applied across categories.
If you have been falsely accused
Do these in order.
Ask for the score and the word count. If the work is under 300 words, the vendor's published reliability figures do not apply to it, and you can say so specifically rather than vaguely.
Ask which tool produced the score and what false positive rate that vendor publishes. Many people running these tools have never seen the number.
Then show your process rather than arguing about the percentage. Version history in Google Docs or Word, intermediate drafts, notes, timestamps, browser history. A probability score cannot be disproved by insisting you did the work, but a document that visibly grew over eleven sessions is very hard to argue with.
That last point matters more than anything else here. Keep your drafts. Not because you are guilty, but because the tools have a known error rate and you may need to be the one who proves it.
Why this is hard to fix
Detectors mostly work by measuring how predictable text is. Machine writing tends to sit in a narrower band of word choice and sentence rhythm than human writing does.
The problem is that some humans naturally write that way. Clear, formal, well-structured prose with conventional phrasing is exactly what a model produces, and it is also what students are taught to produce. Writing advice and detection triggers point in the same direction.
That is why the error will not go to zero, and why the vendor that says so plainly is the one telling you the truth. Every model gets better, then writing tools get better, and the gap narrows again from both sides. If you use AI in your own writing workflow, my how to write a blog post faster with AI piece is about drafting with it rather than hiding it, which is the distinction that matters here.
What I am not going to write
A large share of the search demand around this topic is people looking for ways to beat these tools. Bypass guides, humanizer services, one click removal of AI traces.
I am not writing that, and it is worth saying why rather than quietly leaving it out.
It helps the people least harmed by being caught, since somebody deliberately cheating has already accepted the risk. It does nothing for the person this post is actually for, which is the student flagged for writing their own essay. And it pushes institutions toward scanning harder, which raises the total number of false positives for everyone.
If you use a grammar tool heavily, that is worth knowing about for a different reason. Originality.ai notes that tools which rewrite or edit copy, naming some Grammarly features, can themselves cause false positives in other detectors. My Grammarly review covers what those features do.
The verdict
Two things to take away.
Ask for the false positive rate, not the accuracy rate. Accuracy blends the error nobody pays for with the error somebody pays for, and only one of those ends with a person in a misconduct meeting.
And do the multiplication yourself. At Turnitin's published 1.3 percent, a university running 50,000 submissions a year is producing hundreds of false flags annually, entirely predictably, as designed. Not as a malfunction. That is what the number means, and it is why a detector score should start a conversation rather than finish one.
If you are a student, keep your drafts. If you are a teacher, set a word floor and never let a score be the whole finding. Either way, the tool is a probability, and it has been telling you so in the small print all along. For more on picking tools that hold up, best AI tools for students applies the same standard.
Sourcing
Verified on August 18, 2026 by reading vendor pages directly: GPTZero's 99 percent and 96.5 percent claims and its 1 percent ESL statement, Winston AI's 99.98 percent claim, Originality.ai's 97.8 percent multilingual figure and its false positive guidance including the note about editing tools, Sapling's statement that false positives will occur, ZeroGPT publishing no accuracy figure, and Turnitin's 0.013 and 0.014 false positive rates with the roughly 2,000 sample ELL evaluation and the 300 word threshold. The Writer.com redirect was checked the same day.
Not verified: Copyleaks, Quillbot and Scribbr all return errors to automated requests, and Grammarly's detector page publishes no accuracy figure I could read. Where I could not read a number, I have not given you one. Every accuracy claim in this post is the vendor's own claim about itself, including Turnitin's, and none of these companies test against a shared benchmark.

