Skip to content
Back to BlogIntegrity & AI

How accurate are AI writing detectors, and how often are they wrong?

12 min read2,239 wordsNEW

AI writing detectors return a single score, but that score often gets treated as a verdict, and a false accusation can mean a misconduct hearing or a damaged record. The people most exposed are multilingual writers, whose sentence patterns the tools most often misread. Vanderbilt University examined its own detector and switched it off. This guide gathers the peer-reviewed evidence on how often these tools get it wrong, and what that means if your work is flagged.

Author: MAAS Academic Skills Publishing Desk · Reviewed by a Principal Academic Mentor
Last updated: 2026-08-28
Category: academic-integrity


The honest headline: there is no universal false-positive rate

Direct answer: There is no single accuracy figure for AI writing detectors: the false-positive rate swings widely depending on the tool, the sample and the settings used. False positives, where a detector labels genuine human writing as AI-generated, are real, are not rare, and fall hardest on students who do not write in their first language.

The first thing the research shows is that the studies disagree, and that disagreement is the finding, not a footnote to it. Some evaluations report false-positive rates high enough to make a detector unusable for grading; at least one large study reports zero false positives on its sample. Both can be true at once, because they test different tools, on different kinds of writing, with different thresholds (Elkhatat et al., 2023; Gosling et al., 2024). A detector tuned to almost never miss AI text will flag more human text by mistake; a detector tuned to almost never falsely accuse will let more AI text through. That trade-off is built into how the tools work, so a single reassuring accuracy number, especially one quoted by a vendor, tells you very little about the case in front of you.

This trade-off is why the useful question is not "how accurate is AI detection" in the abstract. It is "how often does this kind of tool get this kind of writer wrong," and there the evidence becomes much sharper.

The number that matters most for multilingual students

Direct answer: The most cited study tested seven GPT detectors against 91 non-native TOEFL essays and 88 US eighth-grader essays. Native-speaker writing was detected accurately, but non-native essays averaged a 61.22% false-positive rate, with 19.78% flagged by all seven detectors at once and 97.8% flagged by at least one.

The most cited study on this question tested seven widely used GPT detectors against 91 TOEFL essays written by non-native English speakers and 88 essays written by US eighth-graders (Liang et al., 2023). For the native-speaker essays, the detectors were close to accurate. For the non-native essays written by real people, the average false-positive rate was 61.22%. Nearly one in five of those essays, 19.78%, was flagged as AI-generated by all seven detectors at once, and 97.8% were flagged by at least one (Liang et al., 2023).

The mechanism explains why this is not a random glitch. Many detectors score a passage on "perplexity," a measure of how surprising or varied its word choices are. Writers working in a second language often use a narrower, more predictable range of vocabulary and sentence structure, which produces low perplexity, which the tool reads as machine-like (Liang et al., 2023). In other words, the feature that flags AI text is also a feature of careful writing in a language you are still mastering. The students least able to absorb a misconduct accusation are the ones the tools are most likely to accuse.

The same problem, one level up: researchers and theses

Direct answer: Giray (2024) found the same bias among scholars worldwide: false positives land disproportionately on non-native English speakers and distinctive writers, and the consequence is a query on a thesis or manuscript, not a capped grade. Researchers who fear being flagged often flatten their own voice, which studies show makes detector scores worse, not better.

The evidence above comes from essays, but the population most exposed to it is not only undergraduates. Giray (2024) examined the experiences of scholars at institutions worldwide and found the same skew: false positives land disproportionately on non-native English speakers and on researchers whose writing style is distinctive. At that level the consequence is not a capped grade. It is a query attached to a thesis, a manuscript, or a research record, judged by a committee or an editor who may have no written procedure for handling a contested detector score.

The exposure is wider than the study-abroad population usually described in this debate. Master's and doctoral programmes across Taiwan, South Korea, Japan and Hong Kong are taught, examined and published in English, so candidates produce all of their formal academic writing in a second language while working in another one daily. Every mechanism in the Liang et al. (2023) results applies to them in full.

Giray (2024) also identifies a cost that accuracy statistics cannot capture. Researchers who fear being flagged begin flattening their own voice, simplifying vocabulary and structure to seem less suspicious. That is precisely the change Liang et al. (2023) showed makes detector scores worse rather than better, so the defensive response feeds the problem it is trying to solve.

When one percent is not small

Direct answer: Turnitin reported about a 1% false-positive rate for its AI detector. Vanderbilt University pointed out that against the roughly 75,000 papers it submitted through Turnitin the previous year, that rate implies about 750 genuine student papers wrongly flagged as AI, which is why Vanderbilt judged the tool unreliable and disabled it.

Institutional data makes the same point at scale. When Turnitin launched its AI detector it reported a false-positive rate of about one percent. Vanderbilt University states that "at the time of launch, Turnitin claimed that its detection tool had a 1% false positive rate" (Vanderbilt University, 2023). Vanderbilt pointed out what that means for a real institution: the university had submitted around 75,000 papers through Turnitin in the previous year, so a one-percent false-positive rate implies roughly 750 pieces of genuine student work wrongly flagged as AI. Vanderbilt concluded the tool was not reliable enough to use for assessment and disabled it (Vanderbilt University, 2023). A "small" percentage becomes hundreds of individual students the moment it meets a real cohort, and each of those students has to prove a negative under pressure.

Vanderbilt was not alone in stepping back from automated AI detection during this period, and its published guidance remains one of the clearest statements of why a low headline error rate can still be unacceptable in practice (Vanderbilt University, 2023).

The UK picture: rising AI use meeting a cohort the tools misread

Direct answer: HEPI's Student Generative AI Survey 2026 found that 94% of UK undergraduates use generative AI for assessed work, and that AI-generated text in submissions rose from 3% in 2024 to 8% in 2025 and 12% in 2026. Rising use raises the pressure to detect, and UK cohorts hold many of the writers these tools misread.

The bias findings above are often read as an American problem, but the conditions that make them consequential are present in the UK system too. HEPI's Student Generative AI Survey 2026 reports that 95% of undergraduates use AI in at least one way and 94% use it to help with assessed work, while the proportion putting AI-generated text into submitted work has risen from 3% in 2024, to 8% in 2025, to 12% in 2026 (Stephenson & Armstrong, 2026). Nearly two-thirds, 65%, say assessment itself has changed significantly in response to AI (Stephenson & Armstrong, 2026). Genuine AI use is climbing, and the institutional pressure to detect it climbs alongside.

Two consequences follow for UK institutions. The first is that the same survey records students voicing anxiety about being falsely accused of misconduct (Stephenson & Armstrong, 2026). That anxiety is the human cost of the false-positive rates documented above, and it registers long before any hearing takes place. The second concerns who is in the room: UK universities teach a large international student population, so the group most likely to be misread by a perplexity-based detector is a substantial share of the cohort rather than an edge case. An error rate that reads as tolerable in the abstract is not distributed evenly across the students who absorb it.

The other side of the evidence

Direct answer: The evidence is not one-sided. One study found a widely used AI detector produced zero false positives across 160 genuine student assignments, vendors have published rebuttals denying bias, and a broader evaluation of several tools found false-positive behaviour was inconsistent rather than uniformly high, with some tools flagging human samples and others not.

An honest read has to hold the contradicting results too, and they are what stop this from being a one-sided complaint. In one study, a widely used generative-AI detector produced zero false positives across 160 genuine student assignments, with every human paper scored at an AI likelihood of zero (Gosling et al., 2024). Detector vendors have also published rebuttals arguing their tools are not biased against non-native writers. And a broader evaluation of several tools found the false-positive behaviour was inconsistent rather than uniformly high: some tools flagged some human samples, others did not (Elkhatat et al., 2023).

Taken together, these results do not cancel out the bias findings; they qualify them. They show that a well-configured tool on a particular sample can perform cleanly, while other tools on other samples, especially samples of non-native writing, perform badly. The responsible conclusion is not "detectors are always wrong." It is that a detector's output is a probabilistic signal with a real and uneven error rate, not proof of anything on its own.

What the evidence adds up to

Direct answer: Three claims survive scrutiny: false positives are real and, on some tools and samples, common rather than marginal; the error is not evenly distributed, with multilingual and non-native writers absorbing far more of it; and the rate is unstable across tools and settings, so a single score should never be treated as a verdict.

Pulling the studies together, three claims survive scrutiny. First, false positives are real and, on some tools and samples, common rather than marginal (Liang et al., 2023; Elkhatat et al., 2023). Second, the error is not evenly distributed: multilingual and non-native writers absorb far more of it, for a reason baked into how the tools measure text (Liang et al., 2023). Third, the rate is unstable across tools and settings, which is precisely why a single score should never be treated as a verdict (Elkhatat et al., 2023; Gosling et al., 2024). This is the same conclusion the institutions that disabled their detectors reached from the other direction (Vanderbilt University, 2023).

The practical upshot is modest and defensible: an AI-detection score is a reason to look more closely, never a reason to conclude on its own that a student cheated. That distinction is the whole difference between a fair process and an unfair one.

What this means if your own work is flagged

Direct answer: If you wrote it and a detector flags it as AI, the evidence above is on your side. An AI-detection score is contested in the research literature, carries a documented false-positive problem, and some universities reject it outright. Your strongest protection is ordinary evidence of your own process: drafts, version history, notes, and the ability to talk through your argument.

If you wrote it and a detector says you did not, the evidence above is your context, and it is on your side. A single AI-detection score is contested in the research literature, carries a documented false-positive problem, and is treated cautiously or rejected outright by universities that examined it closely. Our companion guide on why your own writing gets flagged as AI sets out how the detectors reach that judgment and the concrete steps to respond, and what a similarity score does and does not mean covers the related number students most often confuse with an AI score.

The strongest protection is the ordinary evidence of your own process: drafts, version history, notes, and the ability to talk through your argument. If you are preparing for or responding to a query about your work, an academic integrity check reviews your draft and the evidence trail against the standard your institution applies, and academic support pairs you with a mentor who works with you on your own writing so that the words on the page are demonstrably, and defensibly, yours.

References

Elkhatat, A. M., Elsaid, K., & Almeer, S. (2023). Evaluating the efficacy of AI content detection tools in differentiating between human and AI-generated text. International Journal for Educational Integrity, 19, 17. https://doi.org/10.1007/s40979-023-00140-5

Giray, L. (2024). The problem with false positives: AI detection unfairly accuses scholars of AI plagiarism. The Serials Librarian, 85(5-6), 181–189. https://doi.org/10.1080/0361526X.2024.2433256

Gosling, S. D., Ybarra, K., & Angulo, S. K. (2024). A widely used generative-AI detector yields zero false positives. Aloma, 42(2), 31–43. https://doi.org/10.51698/aloma.2024.42.2.31-43

Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), 100779. https://doi.org/10.1016/j.patter.2023.100779

Stephenson, R., & Armstrong, C. (2026). Student generative artificial intelligence survey 2026 (HEPI Report 199). Higher Education Policy Institute. https://www.hepi.ac.uk/reports/student-generative-ai-survey-2026/

Vanderbilt University. (2023). Guidance on AI detection and why we're disabling Turnitin's AI detector. Brightspace. https://www.vanderbilt.edu/brightspace/2023/08/16/guidance-on-ai-detection-and-why-were-disabling-turnitins-ai-detector/

Share this articleFacebookLinkedInZaloEmail
Want guidance like this?

From this article
to your dissertation.

A 15-minute discovery call: our PhD & Master experts translate this framework into your specific topic and supervisor expectations.