Do AI detectors flag human writing?
Yes, and how often depends on what kind of writing it is and who wrote it. We ran a detector on text written before AI writing was common, so none of it could be AI.
Updated October 1, 2026
How often real writing was flagged
- Random web pages, before 2022
- Texts300
- Flagged as AI1.0%
- Voted forum answers (Stack Exchange), before 2022
- Texts88
- Flagged as AI2.3%
- PAN 2026 human texts (essays, fiction, news)
- Texts1,277
- Flagged as AI3.1%
- Wikipedia and PubMed Central texts (TextSight)
- Texts338
- Flagged as AI3.3%
- Educational web pages, before 2022
- Texts148
- Flagged as AI3.4%
- rasbt human-vs-ai-50k human texts
- Texts2,547
- Flagged as AI6.6%
- LinkedIn posts, 2021
- Texts1,044
- Flagged as AI9.9%
- Literature written before 1940 (TextSight)
- Texts306
- Flagged as AI16.0%
- Wikipedia article openings, before 2022
- Texts64
- Flagged as AI26.6%
| Human writing | Texts | Flagged as AI |
|---|---|---|
| Random web pages, before 2022 | 300 | 1.0% |
| Voted forum answers (Stack Exchange), before 2022 | 88 | 2.3% |
| PAN 2026 human texts (essays, fiction, news) | 1,277 | 3.1% |
| Wikipedia and PubMed Central texts (TextSight) | 338 | 3.3% |
| Educational web pages, before 2022 | 148 | 3.4% |
| rasbt human-vs-ai-50k human texts | 2,547 | 6.6% |
| LinkedIn posts, 2021 | 1,044 | 9.9% |
| Literature written before 1940 (TextSight) | 306 | 16.0% |
| Wikipedia article openings, before 2022 | 64 | 26.6% |
Detector: MELD, at the line its makers set so that 1% of human writing is flagged. Sources: report.txt, report_v2.txt, report_v3.txt, public_report.txt and humanizers_report.txt. On random web text it was right about that. On polished, formal or old-fashioned writing it was far off, and books written before 1940 can't have come from a chatbot.
Why polished writing gets flagged
Detectors learn what AI text looks like, and AI text is tidy, even and confident. So is a Wikipedia opening, and so is a LinkedIn post written to impress. The more a genre rewards a smooth, neutral voice, the more of its human writing crosses the line.
The fix is to set the line per genre. When we set it from 1,044 real LinkedIn posts, the share of real posts flagged dropped from 9.9% to 0.7%, while still catching about half of the AI drafts at that line and about 90% at a looser one. One caveat: the posts we checked that on came from the same 12 authors, so it's the line agreeing with itself, not a fresh test.
Non-native English writers are flagged more
A Stanford study ran seven detectors on 91 TOEFL essays written by non-native English speakers and 88 essays by US eighth-graders. The detectors were near-perfect on the US essays but flagged the TOEFL essays as AI 61.22% of the time on average. All seven agreed on 18 of the 91, and 89 of the 91 were flagged by at least one (Liang et al., 2023).
The authors traced it to perplexity: a writer with a smaller working vocabulary makes more predictable word choices, which is exactly what perplexity-based detectors read as machine-like. When ChatGPT was used to enrich the essays' word choice, the average false-positive rate fell to 11.77%. The essays hadn't become more human; they'd become less predictable.
We haven't measured this with MELD, which is a trained classifier rather than a pure perplexity score, so we can't say how much of the gap it shares. The general lesson holds for any detector: plain, careful, predictable writing looks more like AI.
Even detector makers say it
OpenAI released its own AI-text classifier in January 2023 and withdrew it on July 20, 2023, citing its low rate of accuracy. A detector's score is a probability from a model, not a record of how the text was made.
If you've been flagged and didn't use AI
- 01Ask which detector was used and at what threshold. A score with no false-positive rate behind it isn't evidence.
- 02Show your process: version history in Google Docs or Word, drafts, notes, sources you opened, and earlier writing in the same voice.
- 03Point to the genre problem. Formal, encyclopedic or corporate writing is flagged far more often than casual writing, as the table above shows.
- 04If English isn't your first language, say so and point to the research on non-native writers.
- 05Offer to talk through the piece or write a short section in person. Knowing your own argument is the strongest evidence there is.
Questions
How accurate are AI detectors?+
It depends on the detector and the kind of writing. The one we tested flagged 1% of random human web pages but 27% of human-written Wikipedia openings. Accuracy on one genre doesn't carry over to another.
Why was my own writing flagged as AI?+
Polished, even, neutral writing looks like AI to a detector. Encyclopedic and corporate styles are flagged most, and so is writing with a smaller, more predictable vocabulary. A flag isn't proof of anything on its own.
Are AI detectors biased against non-native English speakers?+
Research says some are. In a 2023 Stanford study, seven detectors flagged non-native writers' TOEFL essays as AI 61.22% of the time on average, while getting US students' essays nearly all right.
Can old writing be flagged as AI?+
Yes. MELD flagged 16% of literature written before 1940 and 27% of Wikipedia openings written before 2022. None of it could have come from a chatbot.
What is a false positive in AI detection?+
Human writing that a detector labels as AI. Detectors pick a line that trades false positives against missed AI; a 1% line means about 1 in 100 human texts of the kind it was tuned on gets flagged, and more on other kinds.
Should a teacher fail a student on a detector score alone?+
We'd say no. Detectors misfire on whole genres and on non-native writers, and OpenAI withdrew its own classifier over low accuracy. Combine any score with the student's drafts, version history and a conversation about the work.
Sources and dates
- MELD, the open-source detector we test with (Hugging Face)huggingface.co
- Liang et al., GPT detectors are biased against non-native English writers (arXiv, Apr 2023; Patterns, Jul 2023)arxiv.org
- rasbt/human-vs-ai-50k dataset (Hugging Face)huggingface.co
- PAN 2026 and RAID texts (Hugging Face)huggingface.co
- TextSight AI text detection benchmark 2026 (Hugging Face)huggingface.co
- Search Engine Journal on OpenAI withdrawing its AI classifier (Jul 25, 2023)searchenginejournal.com
Facts about other products come only from these pages, on the dates shown; prices are as published that day. Macaron's own details are as shipped on October 1, 2026. Products change, so check their sites for the latest.