
The Number Is a Share of Flagged Prose, Not a Verdict
An AI detector percentage reports how much of the qualifying text a model flagged. It does not report how much of your paper a machine wrote, and it says nothing about intent.
Turnitin scores a submission only when it holds at least 300 words of prose, stops at 30,000 words, and prints an asterisk rather than a figure whenever the result lands under 20%. The published document false positive rate sits below 1% above that line, while the sentence level rate runs near 4%.
So read the number as a flag that opens a conversation. A student, a teacher, and an editor each need a different next step, and the last three sections spell those out.
Almost every argument about these tools starts from a misreading of one figure. Someone sees 34% and hears the claim that a third of the essay came from a chatbot.
That is not what the model reported. The percentage is a share of a specific slice of text, produced by a classifier tuned to miss some machine writing rather than accuse a human.
Once the denominator becomes clear, most of the panic drains out of the score. What remains is a narrow, useful signal that belongs beside other evidence rather than above it.
Turnitin Needs 300 Words Before It Scores Anything
The indicator ignores plenty of what sits in your file. Turnitin requires at least 300 words of prose in a long form document, and it rejects submissions over 30,000 words.
Prose is the operative word. Bullet lists, tables, reference entries, and code blocks fall outside the scored text, so the denominator is smaller than the document you uploaded.
That explains a common complaint. A 200-word discussion post comes back with no score at all, and students read the blank result as either an acquittal or a glitch when it is neither.
It also explains why two papers of similar length produce very different figures. A lab report thick with tables offers the model far less qualifying prose than a literature review of the same page count.
The practical move is to ask what the percentage was calculated over. In an appeal, the size of the scored slice matters as much as the number attached to it.
Why a Score Under Twenty Percent Arrives With an Asterisk
- ● The number covers qualifying prose only
- ● Below 20 percent Turnitin shows an asterisk
- ● Sentence highlights carry the higher error rate
Turnitin publishes an unusual admission for a commercial product. Between 0 and 19 percent the tool sees a higher incidence of false positives, so it shows an asterisk instead of a number and withholds the sentence highlights.
Read that band as declining to answer. A paper carrying an asterisk has not been cleared, and it has not been accused either.
The 20% threshold marks where the company stands behind its own figure. Below it, the model has too little signal to separate a careful human writer from a light machine draft.
Anyone building a classroom policy should mirror that line. Treating a 7% result as evidence uses the tool in a range where its maker refuses to.
The Error Rate Lives at the Sentence, Not the Document
Two error rates travel with every report, and only one of them usually gets quoted. Turnitin states a document level false positive rate under 1% for papers that score 20% or higher.
The sentence level rate is roughly 4%. That figure governs the highlighted lines, which are the part a student actually argues about.
Work the arithmetic on a real paper. In a hundred sentence essay, a handful of highlighted lines could be wrong even when the headline percentage behaves exactly as designed.
The company also tunes the balance on purpose. Turnitin reports catching around 85% of unmodified AI text, accepting misses so that false accusations stay rare, which is the right trade for a tool aimed at students.
That design choice carries a consequence people rarely follow through. A low score is weak evidence of human authorship, because the model lets borderline machine text pass.
Two Detectors, Two Different Meanings for One Number
Two tools can print 88 and mean unrelated things. Turnitin reports a share of flagged prose, while GPTZero reports how often its classifier is correct on similar documents.
| What you want to know | Turnitin AI writing indicator | GPTZero |
|---|---|---|
| What the headline number means | Share of qualifying prose flagged, 0 to 100% | Chance the classification is right for that text |
| Minimum text it will score | 300 words of prose | No stated minimum, accuracy rises with length |
| Ceiling per submission | 30,000 words | 150,000 characters on the free tier |
| Low confidence handling | Asterisk instead of a number below 20% | Separate confidence bands reported with the result |
| Stated accuracy | Under 1% document false positives at 20% or above | 99.1% of human articles classed human at high confidence |
| Where the highlights come from | Sentence level, around 4% false positives | Document, then paragraph, then sentence |
The last row carries the practical warning. Both tools grow less reliable as the unit of text shrinks, so a single flagged sentence is the weakest claim either one makes.
Pasting the same paragraph into three free detectors and averaging the answers produces a number with no meaning. The inputs are different scales, and an average of different scales is noise.
What the Percentage Cannot Tell Anyone About You
A detector reads text. It has no access to your drafts, your search history, or the afternoon you spent rewriting a stubborn paragraph.
So the score cannot separate three very different situations that produce similar output. A chatbot draft, a heavily edited human draft, and plain writing by a careful second-language speaker can all land in the same band.
This is why the strongest defense is not a counter score. It is the record of how the document came to exist, a point covered in our guide on what proves you wrote it.
Version history, comments, and dated file copies answer the question the detector cannot. They show a process rather than a probability.
Five Things That Push a Human Paper Upward
Certain writing habits look machine like to a classifier, and none of them involve a chatbot.
Even sentence rhythm. Text with little variation in sentence length matches the smooth statistical profile these models learned to associate with generated prose.
Formulaic transitions. Phrases that open paragraphs the same way each time raise the predictability of the text, which is the exact quality the classifier measures.
Heavy grammar tool editing. Rewriting suggestions flatten idiosyncratic phrasing, and the smoothed result drifts toward the machine end of the scale. Our comparison of Grammarly and QuillBot for students covers where that line sits.
Formal template writing. Lab reports, legal memos, and structured abstracts follow conventions tight enough to read as generated.
Writing in a second language. Learners often reach for safe, high frequency constructions, which is precisely what a predictability model rewards with a higher score.
None of these is misconduct. They are the reason a detector output belongs at the start of a discussion rather than at the end of one.
Who Should Act on a Score and Who Should Let It Go
- ● Ask which sentences were highlighted
- ● Keep version history before you need it
- ● Never compare scores across two tools
The same number calls for different responses depending on who is holding it.
Student with an asterisk result: Let it go, but save your version history now. The tool has declined to make a claim, and your file record is the thing you will wish you had later.
Student flagged above 20%: Ask which sentences were highlighted and request the report rather than the summary figure. Sentence highlights carry a roughly 4% error rate, and that is the argument worth making.
Instructor setting a policy: Anchor the policy to the 20% line the vendor stands behind, and never use the score alone. A detector output that triggers automatic penalties inherits the sentence level error rate on every paper it touches.
Editor screening freelance work: Treat a high score as a prompt to check sourcing instead of authorship, since fabricated citations are the failure that actually costs you. Our walkthrough on fact checking AI writing covers that check.
Writer testing your own draft: Run it if you like, then ignore anything under the asterisk band. Rewriting clean prose to satisfy a classifier makes the writing worse and the score no more meaningful.
Anyone comparing two tools: Stop. The numbers do not share a scale, and a disagreement between detectors tells you about the tools rather than about the text.
Read the Highlights First, Then the Number
The percentage is the least informative part of the report. It compresses a sentence by sentence judgment into one figure, and the compression is where the meaning disappears.
Open the highlights instead. They show which lines the model reacted to, and those lines are checkable against the drafts and notes that produced them.
Then place the number where it belongs, as one input beside the writing process itself. Detectors keep changing, so confirm current thresholds and stated accuracy on each vendor site before you rely on a figure in a formal decision.
For the neighbouring question of how these tools differ from the older kind of check, see our comparison of AI detectors and plagiarism checkers. The two answer different questions, and mixing them up causes as much trouble as misreading the percentage itself.
FAQ
Does a 20 percent AI score mean I will fail the assignment?
Not on its own. The number reports how much qualifying prose a model flagged, and Turnitin states plainly that the score exists to open a conversation. Most institutions treat it as one signal beside draft history, version files, and the writing itself.
What does the asterisk on a Turnitin AI score mean?
It marks a result the tool does not trust. Turnitin found a higher rate of false positives between 0 and 19 percent, so it shows an asterisk instead of a figure and withholds sentence highlights in that band.
How many words does a paper need before Turnitin gives an AI score?
A submission needs at least 300 words of prose before the AI writing indicator runs, and the file caps out at 30,000 words. A 200-word discussion post never receives a score, which is why short assignments come back blank rather than clean.
Is the AI writing score the same as the similarity score?
No. The similarity score measures overlap with sources in the database, while the AI writing indicator estimates authorship of the prose. The two run separately and neither figure feeds the other.
Do free AI detectors give the same percentage as Turnitin?
Rarely. Each tool defines its number differently, so a GPTZero probability and a Turnitin percentage are not measuring the same quantity. Comparing them as if they shared a scale is the most common mistake in an appeal.
Why did my own writing get flagged as AI?
Plain, evenly structured prose. Short declarative sentences, predictable transitions, and heavily edited drafts all shift text toward patterns the models score as machine written, which is why polished second-language writing draws flags more often.
Sources
- GPTZero Support: interpreting results, confidence scores and mixed results — checked 2026-09-27
Some links may be affiliate links. We may earn a commission at no extra cost to you.
This article was written with AI assistance. It is researched and fact-checked, not based on personal hands-on testing unless explicitly stated.
Comments
Post a Comment