Skip to main content

9 July 2026

AI Detection

Do AI Detectors Actually Work? An Honest Answer

AI detectors promise to catch machine-written text, but the false positives, bias, and easy workarounds tell a different story. Here is what actually holds up.

Do AI Detectors Actually Work? An Honest Answer, AI Detection, AI Content analysis by Amjid Ali.

Every few weeks someone forwards me a screenshot of a detector score, panicking that their writing, or their kid’s essay, got flagged as machine written. My answer disappoints them, because it is not the clean yes or no they wanted.

The honest answer is that AI detectors do not work reliably enough to trust as proof of anything. They are a weak signal wearing the costume of hard evidence, and the gap between what they claim and what they deliver has real consequences for real people. Let me walk through what the research actually shows, then what to do instead.

How do AI detectors claim to work?

An AI detector reads a block of text and returns a probability that a machine wrote it. Under the hood, most tools measure two things: perplexity (how predictable the next word is) and burstiness (how much sentence length and rhythm vary). Human writing tends to be less predictable and more varied. Machine writing tends to be smoother and more uniform. The detector looks for that smoothness and scores accordingly.

That is a reasonable idea. The problem is everything that happens once you apply it to the messy reality of how people actually write.

A detector is not reading for meaning, plagiarism, or intent. It is guessing at a statistical fingerprint, and plenty of humans share that fingerprint.

Are AI detectors accurate?

On clean, unedited AI text the good ones are decent, but accuracy falls apart the moment a human touches the text. The vendors quote impressive numbers. Turnitin advertises around 98 percent accuracy with a false positive rate under 1 percent. Independent testing rarely reproduces that.

The bigger issue is what those percentages hide. A 1 percent false positive rate sounds tiny until you remember a single university runs millions of submissions a year. One percent of millions is thousands of innocent people flagged. And the false positive rate is not evenly spread. It lands hardest on the people least able to defend themselves.

Here is how the major tools stack up.

DetectorClaimed accuracyKnown issues
Turnitin AI~98%, <1% false positiveIndependent tests show higher false positives; multiple universities have switched it off
GPTZero~99% on raw AI textFalse positive rates reported above 20% on creative and academic human writing
Originality.aiUp to 100% on raw AI text in vendor testsStruggles with edited and humanised text; elevated false positives for non-native writers
Copyleaks~99% marketedAccuracy drops sharply on paraphrased and mixed content

Notice the pattern. Every tool looks brilliant on raw, untouched machine text, which is exactly the scenario that almost never occurs in the wild. Real submissions are edited, mixed, translated, and rewritten. That is where the numbers quietly collapse.

Can AI detectors be wrong?

Yes, in both directions, and that is the whole problem. There are two failure modes and detectors suffer from both.

False positives flag human writing as AI. This is the dangerous one, because it accuses innocent people. Clear, plain, well-structured writing is the most likely to trip the alarm, which means the better you write to a formula, the more suspicious you look.

False negatives let AI writing pass as human. A short prompt asking the model to “write with varied sentence length and a casual voice” is often enough. Beyond that, an entire industry of humaniser tools exists to launder machine text past detectors. Academic research has shown just how fragile the defence is: one well-known study found that running AI text through the DIPPER paraphraser dropped a leading detector’s accuracy from 70.3 percent to 4.6 percent without meaningfully changing the meaning (Krishna et al.). Read that again. A single paraphrasing pass took detection from usable to worthless.

So the tool punishes honest writers who happen to write cleanly, and waves through anyone determined enough to run their text through a free rewriter. That is close to the opposite of what you would want.

Are AI detectors biased against non-native English writers?

This is the best documented failure, and it should end the debate about using detectors as evidence on its own. A Stanford study ran 91 essays written by real non-native English speakers (genuine human authors) through seven leading detectors. The tools misclassified more than 61 percent of them as AI generated, while flagging native-speaker essays almost never (Liang et al.). One in five essays was unanimously called machine written by every detector tested.

The reason is simple and damning. Second-language writers tend to use simpler vocabulary and more predictable sentence structures. That is precisely the statistical fingerprint detectors read as “machine”. As one summary of the research put it, the tools “cannot tell the two apart” (Tech & Learning).

I take this personally as someone who works across a multicultural profession. A tool that systematically accuses migrants, international students, and ESL professionals of cheating because of how their English is shaped is not a neutral instrument. It is a bias engine with a confidence score.

What happens when a detector gets it wrong?

The consequences are not hypothetical, and they have already reached the courts. Students have been handed failing grades, suspensions, and misconduct findings on the strength of a probability score. Reported cases include a French-born Yale graduate student who says he was wrongly flagged by GPTZero, coerced into a confession, and suspended for a year, and a University of Minnesota doctoral student expelled after an AI accusation who is now suing for denial of due process (Crowell & Moring).

Think about the asymmetry. The detector spends a fraction of a second producing a number. The person on the other end can lose a year of their life, a qualification, or a reputation trying to prove a negative. You cannot easily prove you did not use AI. That burden falls on the accused, and a probability score is a terrible thing to build a life-altering decision on.

Why are universities turning Turnitin’s AI detector off?

Because the institutions closest to the data decided the risk was not worth it. This is not fringe. Through late 2025 and into 2026 a wave of universities disabled AI detection while keeping traditional plagiarism matching. In Australia, Curtin University announced in early 2026 that it would switch off Turnitin’s AI detection across all campuses, explicitly citing reliability and bias concerns (EdTech Innovation Hub). Australian Catholic University, the University of Waterloo, the University of Cape Town, and others made similar calls, most citing the same short list: false positives, equity concerns, and the legal exposure of disciplining someone over a guess.

When the people paying for the tool start turning it off, that tells you something the marketing will not.

To be fair, detection is not useless in every context. As one directional signal among several, reviewed by a human who understands the limits, a score can prompt a closer look. The failure is treating it as a verdict.

What should you do instead?

Stop chasing a magic detector and redesign the way you check work. For educators, the durable answers are about process, not scores:

  • Ask for the working, not just the output. Draft history, version logs, and outlines are far harder to fake than a finished essay and far more revealing.
  • Use short oral checks. A five-minute conversation where a student explains their reasoning tells you more than any percentage. Someone who wrote the work can defend it.
  • Redesign the assessment. Anchor tasks in personal experience, live class discussion, local data, or in-class writing that is genuinely hard to outsource.

For business, the framing is different but the principle holds. Do not police your team or your suppliers with a detector. Set a disclosure policy (say when and how AI was used), and hold the work to a quality and accuracy standard that does not care who or what drafted it. I have written more on operating in an AI-shaped world in my primer on what AI actually means for a business, and on the tools worth adopting in my 2026 Australian AI tools guide. The goal is good work, verified properly, not a witch hunt run by a black box.

There is also a discoverability angle worth naming. As search shifts toward AI-generated answers, the winning strategy is clearly attributed, genuinely useful content, not text tortured to fool a classifier. That is a topic in itself, which I cover in my guide to answer engine optimisation.

The bottom line: AI detectors are a probability dressed up as proof. Use them, if at all, as one soft signal among many, never as evidence, and never as the thing that decides someone’s grade, job, or reputation.

Amjid Ali is an AI and technology leader based in Melbourne who helps organisations adopt AI without losing their judgement. If you are wrestling with AI policy, disclosure, or assessment design, get in touch.

Frequently asked.

Do AI detectors actually work for catching ChatGPT-written content?
Not reliably. On raw, unedited AI text most detectors score above 90 percent, but accuracy collapses on edited, mixed, or paraphrased writing. They also flag genuine human work, so a positive result is a signal to look closer, never proof. Treat any single score as a probability, not a verdict.
Why do AI detectors flag human writing as AI generated?
Detectors measure statistical patterns like predictable word choice and even sentence rhythm. Clear, formulaic, or simply written human text shares those same patterns, so it gets flagged. This is why non-native English writers, students, and technical authors who write plainly are caught far more often than the marketing suggests.
Are AI detectors biased against non-native English speakers?
Yes, and this is the best documented problem. A Stanford study found seven leading detectors wrongly flagged more than 61 percent of essays written by real non-native English speakers as AI generated, while almost never misclassifying native writers. The predictable structure of second-language writing pattern matches what detectors treat as machine text.
Can students get in trouble for a false positive AI detection?
They can, and several have. Real students have faced failing grades, suspensions, and lawsuits after being flagged by tools like Turnitin and GPTZero for work they wrote themselves. Responsible institutions now forbid using a detector score as the sole basis for an academic misconduct finding, and some have switched the tools off entirely.
How can teachers detect AI writing without relying on AI detectors?
Focus on process, not a score. Ask for drafts and version history, hold short oral checks where a student explains their reasoning, and design assessments around personal experience, live data, or in-class work that is hard to outsource. These methods are slower than a button, but they hold up to appeal in a way a probability score never will.

Read another.