Skip to content
Back to BlogWriting Tips

How accurate are AI writing detectors, and how often are they wrong?

7 min read1,323 wordsNEW

There is no single accuracy figure, and anyone who quotes you one is oversimplifying: the false-positive rate of AI writing detectors swings widely depending on the tool, the sample, and the settings.

There is no single accuracy figure, and anyone who quotes you one is oversimplifying: the false-positive rate of AI writing detectors swings widely depending on the tool, the sample, and the settings. What the published evidence does agree on is narrower and more important. False positives, where a detector labels genuine human writing as AI-generated, are real, they are not rare, and they fall hardest on students who do not write in their first language. The problem is serious enough that Vanderbilt University examined the tool and switched its own detector off. This guide pulls the peer-reviewed numbers together in one place, keeps the contradicting studies in view, and explains what the evidence means if your own work is the piece that gets flagged.

Author: MAAS Academic Skills Publishing Desk · Reviewed by a Principal Academic Mentor
Last updated: 2026-07-23
Category: writing-tips


The honest headline: there is no universal false-positive rate

The first thing the research shows is that the studies disagree, and that disagreement is the finding, not a footnote to it. Some evaluations report false-positive rates high enough to make a detector unusable for grading; at least one large study reports zero false positives on its sample. Both can be true at once, because they test different tools, on different kinds of writing, with different thresholds (Elkhatat et al., 2023; Gosling et al., 2024). A detector tuned to almost never miss AI text will flag more human text by mistake; a detector tuned to almost never falsely accuse will let more AI text through. That trade-off is built into how the tools work, so a single reassuring accuracy number, especially one quoted by a vendor, tells you very little about the case in front of you.

This is why the useful question is not "how accurate is AI detection" in the abstract. It is "how often does this kind of tool get this kind of writer wrong," and there the evidence becomes much sharper.

The number that matters most for multilingual students

The most cited study on this question tested seven widely used GPT detectors against 91 TOEFL essays written by non-native English speakers and 88 essays written by US eighth-graders (Liang et al., 2023). For the native-speaker essays, the detectors were close to accurate. For the non-native essays written by real people, the average false-positive rate was 61.22%. Nearly one in five of those essays, 19.78%, was flagged as AI-generated by all seven detectors at once, and 97.8% were flagged by at least one (Liang et al., 2023).

The mechanism explains why this is not a random glitch. Many detectors score a passage on "perplexity," a measure of how surprising or varied its word choices are. Writers working in a second language often use a narrower, more predictable range of vocabulary and sentence structure, which produces low perplexity, which the tool reads as machine-like (Liang et al., 2023). In other words, the feature that flags AI text is also a feature of careful writing in a language you are still mastering. The students least able to absorb a misconduct accusation are the ones the tools are most likely to accuse.

When one percent is not small

Institutional data makes the same point at scale. When Turnitin launched its AI detector it reported a false-positive rate of about one percent. Vanderbilt University pointed out what that means for a real institution: the university had submitted around 75,000 papers through Turnitin in the previous year, so a one-percent false-positive rate implies roughly 750 pieces of genuine student work wrongly flagged as AI. Vanderbilt concluded the tool was not reliable enough to use for assessment and disabled it (Vanderbilt University, 2023). A "small" percentage becomes hundreds of individual students the moment it meets a real cohort, and each of those students has to prove a negative under pressure.

Vanderbilt was not alone in stepping back from automated AI detection during this period, and its published guidance remains one of the clearest statements of why a low headline error rate can still be unacceptable in practice (Vanderbilt University, 2023).

The other side of the evidence

An honest read has to hold the contradicting results too, and they are what stop this from being a one-sided complaint. In one study, a widely used generative-AI detector produced zero false positives across 160 genuine student assignments, with every human paper scored at an AI likelihood of zero (Gosling et al., 2024). Detector vendors have also published rebuttals arguing their tools are not biased against non-native writers. And a broader evaluation of several tools found the false-positive behaviour was inconsistent rather than uniformly high: some tools flagged some human samples, others did not (Elkhatat et al., 2023).

Taken together, these results do not cancel out the bias findings; they qualify them. They show that a well-configured tool on a particular sample can perform cleanly, while other tools on other samples, especially samples of non-native writing, perform badly. The responsible conclusion is not "detectors are always wrong." It is that a detector's output is a probabilistic signal with a real and uneven error rate, not proof of anything on its own.

What the evidence adds up to

Pulling the studies together, three claims survive scrutiny. First, false positives are real and, on some tools and samples, common rather than marginal (Liang et al., 2023; Elkhatat et al., 2023). Second, the error is not evenly distributed: multilingual and non-native writers absorb far more of it, for a reason baked into how the tools measure text (Liang et al., 2023). Third, the rate is unstable across tools and settings, which is precisely why a single score should never be treated as a verdict (Elkhatat et al., 2023; Gosling et al., 2024). This is the same conclusion the institutions that disabled their detectors reached from the other direction (Vanderbilt University, 2023).

The practical upshot is modest and defensible: an AI-detection score is a reason to look more closely, never a reason to conclude on its own that a student cheated. That distinction is the whole difference between a fair process and an unfair one.

What this means if your own work is flagged

If you wrote it and a detector says you did not, the evidence above is your context, and it is on your side. A single AI-detection score is contested in the research literature, carries a documented false-positive problem, and is treated cautiously or rejected outright by universities that examined it closely. Our companion guide on why your own writing gets flagged as AI sets out how the detectors reach that judgment and the concrete steps to respond, and what a similarity score does and does not mean covers the related number students most often confuse with an AI score.

The strongest protection is the ordinary evidence of your own process: drafts, version history, notes, and the ability to talk through your argument. If you are preparing for or responding to a query about your work, an academic integrity check reviews your draft and the evidence trail against the standard your institution applies, and academic support pairs you with a mentor who works with you on your own writing so that the words on the page are demonstrably, and defensibly, yours.

References

Elkhatat, A. M., Elsaid, K., & Almeer, S. (2023). Evaluating the efficacy of AI content detection tools in differentiating between human and AI-generated text. International Journal for Educational Integrity, 19, 17. https://doi.org/10.1007/s40979-023-00140-5

Gosling, S. D., Ybarra, K., & Angulo, S. K. (2024). A widely used generative-AI detector yields zero false positives. Aloma, 42(2), 31–43. https://doi.org/10.51698/aloma.2024.42.2.31-43

Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), 100779. https://doi.org/10.1016/j.patter.2023.100779

Vanderbilt University. (2023, August 16). Guidance on AI detection and why we're disabling Turnitin's AI detector. Brightspace. https://www.vanderbilt.edu/brightspace/2023/08/16/guidance-on-ai-detection-and-why-were-disabling-turnitins-ai-detector/

Share this articleFacebookLinkedInZaloEmail
Want guidance like this?

From this article
to your dissertation.

A 15-minute discovery call: our PhD & Master experts translate this framework into your specific topic and supervisor expectations.