AI writing

How do AI detectors work?

An AI detector reads your text, gives it a score, and calls it AI if the score crosses a line someone chose. There are three main ways to get that score, and each fails in its own way.

Updated October 1, 2026

1. Predictability: perplexity and burstiness

A language model can say how surprised it is by each next word. Averaged over a text, that's perplexity. AI text tends to pick likely words, so it scores low. People make odder choices, so they usually score higher.

Burstiness is how much that varies across a text: people write a short sentence, then a long one, a plain passage, then a strange one. Models tend to keep an even pace.

We measured the rhythm part on our held-out sets. In 300 human explainers and forum answers, sentence lengths in a typical text varied by 8.4 words around their average; in AI drafts of the same pieces, by 6.9. In 300 LinkedIn posts the gap almost disappeared once you allow for AI's shorter sentences. Burstiness is a real signal, but a weak one on its own.

Newer zero-shot methods compare two models' views of the same text. Binoculars, from 2024, reported catching over 90% of ChatGPT text while flagging 0.01% of human text, without training on ChatGPT output (Hans et al.).

The weakness: predictable human writing scores low too. Researchers traced detectors' high false-positive rate on non-native English writers to exactly this (Liang et al., 2023).

2. Trained classifiers

Most detectors you can paste text into are classifiers: a model trained on large piles of human and AI text to tell them apart. It learns the predictability signals above along with many others, like word choice, sentence shapes and how ideas are linked.

MELD, the open-source detector we test with, works this way. It scores the text token by token and pools those into one score for the whole text, which can also be pooled per paragraph to point at the passages that read as AI. On public test sets it separated human from plain AI text almost perfectly (AUROC 0.98–1.00, where 1 is perfect ranking; public_report.txt).

The weakness: a classifier knows what it was trained on. Writing styles it saw little of, like encyclopedia openings or LinkedIn posts, get misread.

3. Watermarks

A watermark is put in at writing time. In one well-known design, before each word the model splits its vocabulary into a random "green" and "red" list and leans toward green words. Anyone with the key can count the green words and run a statistical test: human writing lands on green only as often as chance predicts, watermarked text far more often (Kirchenbauer et al., 2023).

Google's SynthID Text works in a similar spirit and is open source. Its documentation says the watermark is less effective on factual answers and that detection confidence can drop sharply when text is thoroughly rewritten or translated.

The weakness: a watermark only exists if the model that wrote the text added one, and only the key holder can check it. A detector you paste text into is usually not reading a watermark.

The line

Whatever the method, the score only becomes a verdict when it's compared with a threshold. Thresholds are usually set by false-positive rate: the line where 1 in 100, or 1 in 20, real human texts get flagged. A stricter line flags less human writing and misses more AI.

1 in 100
Real posts flagged0.7%
AI drafts caught50–57%
1 in 20
Real posts flagged5.3%
AI drafts caught88–95%

Round 3, report_v3.txt: 300 held-out posts, drafts from Gemma 4 and Qwen3.5. Any detector that claims to catch everything is flagging a lot of people too.

What detectors are bad at

  • Short text. Under about 40 words the score is mostly noise, so we never judge a piece that short on its own. OpenAI said its own classifier, withdrawn in July 2023, was unreliable under 1,000 characters.
  • Other genres. A line set on web pages flagged 27% of human Wikipedia openings and 16% of pre-1940 literature.
  • Naming the model. Our detector called Gemma and Qwen drafts GPT or Claude 96% of the time. Models trained on similar data write alike.
  • Heavily rewritten text. A sentence-level rewrite took our pilot drafts from 93.3% flagged to 19.3% in one try.

Questions

What do AI detectors look for?+

How predictable the word choices are, how even the rhythm is, and, in trained detectors, many patterns learned from examples of human and AI text. Visible habits like em dashes are a small part of it.

What are perplexity and burstiness?+

Perplexity is how surprised a language model is by each next word; AI text is usually less surprising. Burstiness is how much that, and sentence length, varies across a text; people vary more. In our data the rhythm gap was clear in explainers and small in LinkedIn posts.

Can AI detectors tell which AI wrote something?+

Not reliably. The detector we use guessed the model family for 600 AI drafts and called them GPT or Claude 96% of the time; they were all from Gemma and Qwen.

Do ChatGPT and Gemini watermark their text?+

Google says it has open-sourced SynthID Text, its text watermark. Whether a given chatbot's text carries a watermark is up to its maker, and only the key holder can check for it.

Are AI detectors reliable?+

As a signal, sometimes; as proof, no. Every detector trades missed AI against wrongly flagged people, and the balance shifts with the kind of writing. OpenAI withdrew its own classifier in 2023, citing low accuracy.

How long does text need to be for a detector?+

Longer is better. We don't judge anything under about 40 words on its own: there's too little text for the score to mean much.

Sources and dates

Facts about other products come only from these pages, on the dates shown; prices are as published that day. Macaron's own details are as shipped on October 1, 2026. Products change, so check their sites for the latest.

Keep reading