
Two Words for a Method the Leading Detector Retired
- ● Perplexity - how predictable your words were
- ● Burstiness - how much that varies per sentence
- ● Retired as the engine in autumn 2023
The short answer is that perplexity measures how predictable your wording was to a language model, and burstiness measures how much that predictability swings between sentences. Machine text scored low on both. The awkward part is that GPTZero, the tool that popularised the pair, says it stopped using them for detection in autumn 2023 after moving to a deep learning architecture. They now sit as one of seven indicators rather than as the engine, alongside a classifier the company credits with 99% accuracy in its own testing.
So the vocabulary outlived the technique. Most explainers you will read still describe a method the field has largely moved past, which matters if you are trying to understand a score you were given.
Perplexity Is a Bet on Your Next Word
Perplexity asks a simple question of every sentence. How likely was it that a language model would have picked these exact words in this exact order.
GPTZero illustrates it with a fragment. The phrase beginning I am an AI would probably continue with the word assistant, which is a low perplexity completion, while continuing with the word potato would be high perplexity.
Low perplexity therefore reads as machine written, because a model reaching for the most probable next word produces text that another model finds unsurprising. High perplexity reads as human, because people reach for odd words.
The flaw sits inside that logic. Plain, careful, conventional prose is also unsurprising, so the measure penalises clarity as if it were a symptom.
Burstiness Is the Variation Between Your Sentences
Burstiness looks at the same perplexity scores from one step back. Rather than asking how predictable a sentence was, it asks how much that predictability changes across the document.
The original method calculated it as the standard deviation of the per sentence perplexity scores. A human draft mixes a fourteen word sentence with a three word one, and mixes a plain clause with an odd turn of phrase.
Model output tends to be flatter. Sentences arrive at similar lengths with similar levels of predictability, which produces a low standard deviation and therefore low burstiness.
Unlike perplexity, this measure never had a published cutoff. GPTZero states plainly that there is no set threshold for burstiness and that it was taken as one element among several.
What Low Burstiness Looks Like on the Page
The measure is easier to trust once you can see it. Below are two versions of the same three sentences, written here purely as an illustration of the contrast.
The library opens at nine. The staff are helpful and friendly. The study rooms can be booked online.
Every sentence runs to a similar length, opens with the same article, and lands on a plain predicate. Perplexity would be low across all three, and the variance between them would be lower still.
The library opens at nine. Ask at the desk, though, because the staff know which study rooms are actually quiet. Booking online works, mostly.
The second version carries a fragmentary clause, a hedge at the end, and sentences of noticeably different lengths. That unevenness is what burstiness was built to notice.
Neither passage proves anything about who wrote it. That is the honest limit of the whole approach, since a careful human writer can produce the first pattern on any given day.
The Number 85 and What It Was Ever Meant to Mean
One figure from the original model still circulates widely, so it is worth pinning down. GPTZero wrote that generally a perplexity above 85 is more likely than not from a human source.
Read the hedging in that sentence. It says more likely than not, which describes a lean rather than a finding, and it applied to a model that has since been replaced.
The number also does not convert into anything you see today. Modern detectors report a percentage likelihood across a document, and that percentage is not a rescaled perplexity score.
If someone quotes 85 at you as a standard, they are quoting a retired heuristic from a superseded system. Our comparison of detectors against plagiarism checkers covers the wider habit of treating these tools as if they measure the same thing.
Why a Statistical Method Gave Way to a Trained One
The two measures share a weakness that no amount of tuning fixes. They describe surface predictability, and predictability is a property of good expository writing as much as of generated text.
That produces false positives on exactly the people least able to argue back. Technical documentation, formal academic phrasing, and English written by a non-native speaker all favour common constructions by design.
It also produces false negatives once anyone edits. Breaking up sentences and swapping a few words raises the variance without changing who or what wrote the draft.
A trained classifier is not immune to either problem, but it is not restricted to two summary statistics. It can weigh syntax, word choice, and semantic coherence together rather than reducing a document to a mean and a standard deviation.
What the Current Detectors Weigh Instead
- ● Deep learning classifier replaced the statistics
- ● Internet text search added as a signal
- ● Old metrics kept as 1 of 7 indicators
The table below separates what the original method used from what the current generation leans on. It is the quickest way to see how much of the popular explanation is out of date.
| Signal | What it looks at | Era | Role in the current GPTZero model |
|---|---|---|---|
| Perplexity | How predictable each sentence was to a language model | Original statistical model | One of seven indicators, not the engine |
| Burstiness | The variation in perplexity across sentences | Original statistical model | One of seven indicators, not the engine |
| Deep learning classifier | Learned patterns across syntax, word choice, and coherence | Autumn 2023 onward | The primary detection method |
| Internet text search | Whether the passage already exists online | Added after the rewrite | Supporting signal |
| Sentence level scanning | Which specific sentences look generated | Added after the rewrite | Powers the highlighted output |
| Paraphraser defences | Whether output was reworded to evade detection | Added after the rewrite | Guards against bypass tools |
Details reflect GPTZero published descriptions as of Sep 2026, and detection vendors revise these systems frequently. Confirm the current methodology on the official site before citing any of it in a policy document.
Notice what happened to the two famous metrics. They were not proven wrong so much as demoted, which is a normal outcome for an early heuristic.
Why Watermarking Was Supposed To Settle It
There is a fourth approach that comes up whenever detection accuracy is questioned, and it is worth knowing why it has not taken over. Watermarking embeds a hidden statistical signal in text at the moment a model generates it.
The appeal is obvious. Rather than guessing from the surface of a document, a checker could look for a signal that was deliberately planted, which turns detection from inference into verification.
The problem is that the signal is fragile and voluntary. GPTZero notes that watermarks can be tampered with, and a mark only exists if the model provider chose to add one, which leaves every open model outside the scheme.
Rewriting also erodes it. Paraphrasing tools, translation, and ordinary human editing all disturb the exact token choices a watermark depends on, so the signal weakens precisely when someone is trying to hide the source.
The Accuracy Figures Are Vendor Figures
- ● 99% accuracy, 1% false positive - self reported
- ● 96.5% on mixed human and AI documents
- ● No independent audit behind the headline
Every number you will see quoted for detector accuracy comes from the company selling the detector, and that deserves stating plainly. GPTZero reports 99% accuracy with a 1% false positive rate on AI against human samples, and 96.5% on documents mixing both.
It also cites a third party RAID benchmark identifying 95.7% of AI written text while misclassifying only 1% of human writing. That is a stronger form of evidence than an internal test, though the vendor still chose which benchmark to feature.
Apply the 1% figure to a realistic setting and the problem becomes visible. A class of 300 essays, all genuinely human, still yields roughly 3 accused students at that rate.
The rate is not the risk. The consequence of a single false positive is, which is why detector output belongs in a conversation rather than in an automatic penalty. Our guide to fact checking AI assisted writing sets out the sort of evidence that actually holds up.
What To Do If a Detector Flagged Your Work
A flag is a probability estimate about text, not a finding about a person. The response that works is documentary rather than rhetorical.
Version history is the strongest evidence available to most people. A document with hours of incremental edits, revisions, and abandoned paragraphs is very hard to fake and easy to show.
Explaining your own argument in person is the second. Someone who wrote a piece can say why they chose a structure, which is a question generated text cannot answer for its user.
Arguing about perplexity is the weakest move of the three. You would be debating a metric the leading vendor has already retired, and the person reading your score is not using it either.
Which Reading of a Detector Score Fits Your Situation
A teacher deciding what to do with a flagged essay: treat the score as a prompt to talk, never as a verdict. The vendor own false positive rate makes a single flag an unsafe basis for a penalty.
A student whose honest draft was flagged: collect version history first and ignore the metric vocabulary entirely. Evidence of process beats any argument about how the tool works.
An editor screening freelance submissions: use it as one input beside the writer track record and a short conversation about the piece. Detection scores are least reliable on clean, plain, professional prose.
A marketer worried about search visibility: the score is not a ranking signal, and the question you actually care about is quality and originality, which we cover in whether AI writing tools hurt SEO.
Anyone writing a policy document: name the evidence you will accept rather than a tool or a threshold. Vendors change their methods, and a policy pinned to a retired metric ages badly.
Judge the Claim, Not the Vocabulary
Perplexity and burstiness are worth understanding precisely because they are quoted so confidently by people describing a system that no longer runs on them. Knowing what they measured tells you what the early tools could and could not see.
The durable lesson is not about either metric. It is that predictability and authorship are different things, and every detector built so far has had to guess across that gap.
So ask what a score is actually measuring before you act on it. A number with a retired method behind it is a conversation starter, and it was never designed to be the last word.
FAQ
What do perplexity and burstiness mean in AI detection?
Perplexity measures how predictable your word choices were to a language model, and burstiness measures how much that predictability varies from sentence to sentence. Text that is both highly predictable and evenly predictable looked machine written under the original method. Both terms describe a statistical approach that the best known detector has since moved away from.
Do AI detectors still use perplexity and burstiness?
GPTZero states that as of autumn 2023 it no longer uses perplexity and burstiness for detection, having migrated to a deep learning architecture. The two measures survive as one of seven indicators in the newer model rather than as the engine. Many smaller tools and most explainer articles still describe the retired method.
What does a perplexity score above 85 mean?
It was the rough dividing line in the original statistical model, where GPTZero said a perplexity above 85 was more likely than not from a human source. It was never a legal or academic standard, and it does not map onto the percentage scores detectors show today. Treat any single cutoff you read about as a historical detail.
Why does a detector flag writing that a person actually wrote?
Because clear, conventional prose is predictable by design. Technical writing, careful academic phrasing, and English written by a non-native speaker all tend to use common constructions, which is exactly what a predictability score rewards. That is the structural reason the original method produced false positives on real human work.
How accurate are AI detectors really?
The published figures come from the vendor. GPTZero cites 99% accuracy with a 1% false positive rate on clean AI against human samples, and 96.5% on mixed documents. Those are self-reported benchmark results rather than independent audits, so read them as a claim about ideal conditions rather than a guarantee about your document.
Sources
- GPTZero: Perplexity, burstiness, and statistical AI detection — checked 2026-09-27
- GPTZero Support: interpreting results, confidence scores and mixed results — checked 2026-09-27
Some links may be affiliate links. We may earn a commission at no extra cost to you.
This article was written with AI assistance. It is researched and fact-checked, not based on personal hands-on testing unless explicitly stated.
Comments
Post a Comment