Are AI content detectors accurate? What the evidence says
A published test flagged most essays by non-native English writers as AI. OpenAI withdrew its own detector. Here is what that means for judging a web page.
AI content detectors promise to say whether a piece of writing came from a chatbot. The published evidence says they are not reliable enough to settle that question for a single page, and they fail unevenly. This guide covers what was tested, what OpenAI did with its own detector, and what to look at instead when you are judging a website. We make Kitsuvo, a browser that detects AI-built sites, and we explain at the end why it does not lean on a text detector.
What a 2023 study found
Researchers tested seven widely used GPT detectors in a paper published in Patterns on July 10, 2023, GPT detectors are biased against non-native English writers. They ran 91 TOEFL essays written by non-native English speakers and 88 essays by US eighth-graders through the detectors. The detectors classified the US essays accurately. On the TOEFL essays they reported an average false-positive rate of 61.3%, meaning human writing was labelled AI. At least one detector flagged 97.8% of the TOEFL essays as AI-generated.
The authors link the result to perplexity, a measure of how predictable the word choices are. Most GPT detectors use it, and writers with a smaller range of expressions score as more predictable, so their work looks machine-made. This is a 2023 test of detectors available then, so it describes a failure mode rather than any one tool's current accuracy.
Why OpenAI withdrew its own detector
OpenAI discontinued its AI text classifier in July 2023, citing its low rate of accuracy, as The Decoder reported. In its tests, per that report, the tool labelled human-written text as AI-written about 9% of the time. If the company behind ChatGPT did not keep a detector running, that is a reason for caution with a third-party score.
Google does not need a detector
Search engines face the same question at web scale, and Google's answer is not to guess at authorship. Its guidance on generative AI content says AI can help with research and structure, and that using it to generate many pages without adding value for users may violate its spam policy on scaled content abuse. The test is value and scale, not who or what typed the words. The same logic is why a polished page and a mass-produced one can both come from a chatbot, and only one is a problem.
What works better for web pages
A web page leaves evidence that prose does not. Look at how the page was built and published:
- The address and hosting: a site on
lovable.apporbolt.hostwas published through an AI builder. See how to tell if a website was made with AI. - Leftovers: unchanged scaffold titles, favicons and placeholder text, covered in the phrases that give away AI-written website copy.
- The wider site: no named author, no real contact details, dozens of near-identical pages.
Kitsuvo's score is built from signals like these, shown with their evidence, and it labels results as estimates. It can run an optional text and vision estimate through a model you install yourself, with Ollama. Cloud verdicts never decide the filter, and unlabelled or unreadable content can pass, per the 0.2.4 release notes. It does not call a page human-made just because nothing was found. For how web checkers, extensions and text detectors compare, see AI website detector tools compared.
When a detector is still useful
A detector score is one weak signal. It is reasonable as a prompt to read a passage more closely, to compare writing against a person's earlier work, or to check a source. It is not reasonable as proof about a writer, especially one writing in a second language. Treat a high score on a short or formulaic text with the most caution.
Questions people ask
Do AI detectors work on non-native English writing?
Poorly, in the one published test we cite. A 2023 Patterns study found an average false-positive rate of 61.3% across seven detectors on TOEFL essays by non-native English writers, and at least one detector flagged 97.8% of them.
Can an AI detector prove a page was written by AI?
No. A detector returns a probability about the prose. It cannot see who edited the text or whether the writer is a non-native speaker, and OpenAI withdrew its own detector for low accuracy.
Does Google penalize AI-written pages?
Google says using AI is not banned. Its spam policy targets using generative AI to produce many pages that add no value for users, which it calls scaled content abuse.
Does Kitsuvo use an AI text detector?
Not as its main signal. Kitsuvo scores a page from builder fingerprints and other page evidence. An optional local text model can add an estimate, cloud verdicts never decide the filter, and every score is an estimate, not proof.