Best AI Detectors in 2026: What the Independent Studies Say

Best AI Detectors in 2026: What the Independent Studies Say

By David Kim, News & Analysis Editorial Desk · September 29, 2026 · 13 min read

Updated September 29, 2026
Quick Answer

AI detectors are sold on accuracy percentages that independent research does not support in the conditions they are actually used in. The most cited peer-reviewed study, published in Patterns in 2023, ran seven detectors against 91 TOEFL essays written by non-native English speakers and found an average false-positive rate of 61.3 percent, while the same detectors were near-perfect on essays by US eighth graders. Rewriting those essays with more advanced vocabulary cut the misclassification rate from 61.3 percent to 11.6 percent, which shows the tools are partly measuring linguistic simplicity rather than authorship. The RAID benchmark, presented at ACL in 2024, tested detectors across 672,000 texts, 11 domains and 12 adversarial attacks, and found performance collapsing under paraphrasing in some domains. On pricing, Originality.ai is the only major vendor whose full plan structure we could read at source: free with 60 credits a month, Pro at $12.95 a month billed annually or $14.95 monthly for 2,000 credits, and Enterprise at $136.58 annually or $179 monthly, where one credit scans 100 words. GPTZero and Copyleaks render prices client-side and we publish none. Use these tools to start a conversation, never as evidence.

Two true numbers

Vendors in this category advertise accuracy at or above 99 percent. A peer-reviewed study found an average false-positive rate of 61.3 percent.

Both figures are real. They are measuring different things, and the gap between them is the entire subject of this article.

A vendor accuracy number is usually produced by running a detector over clean text from a known model and clean text from a known human, and counting. If you want the mechanism behind the text being detected, our explainer on how generative AI works covers the prediction process these tools are trying to reverse. The 61.3 percent comes from running detectors over essays written by real people who learned English as a second language. Nobody is lying. One test resembles the world and the other does not.

How we assessed these tools

We read pricing from each vendor's own page on 29 September 2026, and we read the independent research rather than the vendors' summaries of it. Where a company's pricing page renders figures client-side and would not return them to an automated read, we publish no price, exactly as we do elsewhere in our tool coverage.

We have not run our own detection tests. We are not in a position to benchmark seven classifiers responsibly, and the published academic work is better than anything we could produce.

What the independent research actually found

The non-native English finding

The most cited paper in this field is GPT detectors are biased against non-native English writers by Liang, Yuksekgonul, Mao, Wu and Zou, published in Patterns in 2023. The researchers ran seven detectors against 91 TOEFL essays written by non-native English speakers and against essays by US eighth-grade students.

ResultFigure
------
Average false-positive rate on TOEFL essays61.3%
TOEFL essays flagged by at least one detector97.8% (89 of 91)
TOEFL essays flagged by all seven detectors19.78% (18 of 91)
Accuracy on US eighth-grade essaysNear-perfect
Misclassification after enriching vocabulary11.6%, down from 61.3%

That last row is the one that explains the mechanism. When the same essays were rewritten with more advanced vocabulary, the false-positive rate fell by a factor of five. The detectors were not identifying authorship. They were identifying simple, predictable language, which is what a statistical model flags and also what a second-language writer often produces.

A correction worth making plainly: this 61.3 percent figure is regularly attributed to Turnitin in particular. It is not a Turnitin measurement. It comes from a 2023 study of seven detectors, none of which is necessarily in the state it was then. Repeating it as a fact about one product today is exactly the kind of number-laundering this category is full of.

The robustness finding

The RAID benchmark, presented at ACL in 2024, is the most serious evaluation infrastructure the field has. It covers 672,000 texts across 11 domains and 12 adversarial attacks including paraphrasing, synonym swaps and deliberate misspellings.

Two things make RAID more useful than any vendor chart. First, it reports results at a fixed false-positive rate, usually 1 percent, which forces detectors to state how many AI texts they catch while wrongly accusing only one human in a hundred. Second, it tests adversarial conditions, and the results move enormously: detectors that look excellent on one domain deteriorate sharply on another or under paraphrasing.

The tools

1. Originality.ai — Best documented pricing, aimed at publishers

Best for: content teams that want detection plus plagiarism and fact-checking in one subscription, with a price they can read.

Originality is the only vendor here whose complete plan structure we could read at source. Basic is free with 60 credits a month, limited to three scans a day and a 2,000-word scan limit. Pro is $12.95 a month billed annually, shown as $155.40 a year, or $14.95 month to month, for 2,000 credits with plagiarism checking, grammar, readability, fact-checking, a Chrome extension and team management. Enterprise is $136.58 a month billed annually, shown as $1,638.96, or $179 monthly, for 15,000 credits, API access and 365-day scan history.

2. GPTZero — Best known in education, price not readable at source

Best for: teachers and institutions, with the caveat that we could not verify what it costs.

GPTZero is the detector most commonly deployed in education, and the only one here that publishes a detailed methodology page. On that page it states a 96.5 percent accuracy rate on mixed documents containing both AI and human writing, a false-positive rate of no more than 1 percent when evaluating AI against human text, and that work since April 2022 reduced its false-positive rate on TOEFL texts to 1.1 percent.

Those are the vendor's own figures. Independent work reported alongside RAID describes GPTZero detecting around 95.7 percent of AI texts at a 1 percent human false-positive rate and being unusually robust to adversarial attacks, which is a genuinely strong result if it holds in your domain.

3. Copyleaks — Enterprise and LMS integrations, price not readable at source

Best for: institutions that need detection embedded in an existing learning platform.

Copyleaks sells primarily to enterprises and education systems, with integrations rather than a consumer workflow, and it appears consistently in independent comparisons as one of the stronger performers on paraphrased text, an area where most detectors struggle badly.

4. Turnitin — the institutional default, and the one being switched off

Best for: institutions that already have it, used with considerable care.

Turnitin is the most widely deployed academic integrity product in the world and its AI detection is sold as part of institutional licensing rather than at a public price, so there is nothing for us to verify.

What is documented is the retreat. Vanderbilt University disabled Turnitin's AI detector and published its reasoning, noting that with 75,000 papers submitted in a year even a 1 percent false-positive rate would mean hundreds of students wrongly flagged. Northwestern disabled AI detection entirely, and reporting in 2026 describes at least a dozen universities doing the same, with several institutions barring detector output as the sole basis for a misconduct case.

Which should you choose?

If you are a publisher or content team: Originality.ai at $12.95 a month annually, used as a triage signal on freelance submissions rather than as a verdict. The credit-to-word conversion is published, which makes the cost predictable.

If you are an educator: GPTZero has done more public work on the bias problem than anyone else in the category, and you will need to get its price from its own page. Our guide to platforms for learning AI covers the other half of this problem, which is teaching students to use these tools openly. Use it to open a conversation, never to close one.

If you need detection inside an existing LMS: Copyleaks, confirming the price and the paraphrase performance in a trial on your own student writing rather than on demo text.

If your institution already runs Turnitin: read Vanderbilt's published reasoning before you rely on the AI score, and check whether your own academic integrity policy still permits it as evidence.

If your writers include a lot of second-language speakers: weight every result far more sceptically, because that is precisely the population the research shows these tools fail on. The same caution applies to our own field: as our guide to AI search visibility tools notes, measurement tools in young categories tend to be sold with more confidence than the evidence supports.

Conclusion

There is a version of this article that ranks five detectors by advertised accuracy and recommends the one with the biggest number. It would be easier to write, and it would be worthless, because the advertised numbers are produced under conditions that do not resemble the essays, articles and applications these tools are pointed at.

The defensible position is narrower. Detection can tell you that a document is worth a second look. It cannot tell you who wrote it, and the published research shows it is most likely to be wrong about people who learned English as a second language. Any process that treats a percentage from a classifier as proof is building a disciplinary decision on a probability, and the institutions with the most experience of doing so are the ones turning it off.

This assessment is an editorial synthesis of vendor pricing pages, vendor methodology statements and peer-reviewed research read on 29 September 2026 and linked inline. We did not run our own detection tests. Where a vendor's prices are rendered client-side and could not be read at source, we publish no figure rather than repeat an unverified one. Vendor accuracy claims are labelled as such throughout. This is not legal advice, and detector output should not be treated as evidence of misconduct.

Key Takeaways

  • The 2023 study in Patterns ran seven GPT detectors against 91 TOEFL essays by non-native English writers and found an average false-positive rate of 61.3 percent; 97.8 percent of those essays were flagged by at least one detector and 19.78 percent by all seven.
  • The same detectors were near-perfect on essays written by US eighth-grade students, which is the finding that matters: the error is not random, it falls on one group of writers.
  • Rewriting the TOEFL essays with richer vocabulary dropped misclassification from 61.3 percent to 11.6 percent, indicating detectors partly measure the simplicity of the writing rather than who produced it.
  • That 61.3 percent figure is frequently misattributed to Turnitin specifically. It comes from a study of seven detectors in 2023 and is not a measurement of any single commercial product today.
  • The RAID benchmark, published at ACL 2024, evaluated detectors across 672,000 texts, 11 domains and 12 adversarial attacks, and reports results at a fixed 1 percent false-positive rate, which is a far more honest measure than a headline accuracy percentage.
  • Originality.ai publishes its full pricing: free with 60 credits a month, $12.95 a month billed annually or $14.95 monthly for 2,000 credits, and $136.58 annually or $179 monthly for 15,000 credits, where one credit covers 100 words.
  • Several universities, including Vanderbilt and Northwestern, disabled Turnitin's AI detection over false-positive concerns, and a number of institutions now bar detector output as sole evidence of misconduct.

Frequently Asked Questions

Are AI detectors accurate?

They are accurate in benchmark conditions and much less so in the conditions they are used in. Independent research finds heavy false positives on non-native English writing and sharp performance drops on paraphrased text, while vendor figures are usually produced on clean, unedited AI output that nobody submits in practice.

Which AI detector is most accurate?

Published evidence does not support a single winner. The RAID benchmark shows results varying dramatically by domain and by adversarial attack, with detectors that lead in one category failing in another. Any tool claiming a single accuracy number across all conditions is describing a test, not your use case.

Can a detector result be used to accuse a student of cheating?

It should not be used alone. Vanderbilt and Northwestern disabled Turnitin's AI detection over false positives, and several universities now prohibit relying on detector output as the sole basis for a misconduct finding. Use it to prompt a conversation about process, drafts and version history instead.

Why do detectors flag non-native English writers so often?

Most detectors score text on statistical predictability. Writing that uses common words and simple constructions looks predictable, and so does machine-generated text. Non-native writers often write in exactly that register, so the tools penalise them for the vocabulary they have rather than for anything they did.

Does paraphrasing defeat AI detectors?

Often, yes. The RAID benchmark includes paraphrasing among its adversarial attacks and shows detector performance deteriorating substantially under it in several domains. Any workflow that assumes a detector will catch deliberately rewritten AI output is assuming something the research does not support.

What should schools and publishers do instead?

Assess process rather than output: drafts, version history, oral defence, in-class writing and source trails. Detection can be one signal among several, but the evidence base does not support treating a percentage from a detector as proof of anything on its own.

About the Author

David Kim avatar

David Kim

News & Analysis Editorial Desk

News & Analysis Editorial Desk · Web3AIBlog

David Kim is a pen name for our news and analysis editorial desk. Posts under this byline are written and reviewed by contributors covering emerging-technology policy, regulatory action, market events, and incident reporting across crypto and AI. The desk emphasizes primary-source reporting (court filings, regulatory text, on-chain data, official postmortems) over reaction-cycle commentary. Every news post links to the underlying source documents so readers can verify the facts.