Predictability
A model can estimate how expected each next word is under a language model. Text with unusually predictable choices may receive a higher machine-likeness score, but polished human prose can also be predictable.
Don't stop here
Hand-picked guides our readers explore right after this one.
Master the 8-step framework for writing prompts that get results
Read the guideExpert guide to Claude prompts with XML tags, artifacts, and complex reasoning
Read the guideAI prompts for marketing campaigns, content creation, SEO, email, and analytics
Read the guideAI literacy and evidence guide · Checked August 13, 2026
A plain-English explanation of classifiers, predictability, writing variation, provenance, false positives, bias, evasion, and a fair review process for schools, employers, and publishers.
Michael Okeje
AI detection and responsible-use research · Last updated August 13, 2026
An AI writing detector does not watch a person write. It receives a piece of text and estimates how closely that text resembles examples associated with machine-generated language. That distinction sounds small, but it changes everything about how the result should be interpreted. A score is an uncertain signal about a document. It is not a fingerprint, a recording, or a proof of who created the words.
The commercial language around detectors often makes the tool sound more certain than the underlying task allows. A dashboard may show “AI probability,” “likely AI,” or a highlighted percentage. Readers understandably convert that display into a statement such as “this student used ChatGPT.” That conversion is not justified unless the tool has been independently validated for the exact language, genre, model generation, sample length, and decision context, and even then the result is usually one piece of evidence rather than a verdict.
I would use detection, when policy permits it, as a triage signal for a process review. I would not use it as an automatic accusation. The person affected deserves to know the rule, see the relevant concern, respond with their process, and receive a decision based on consistent evidence. The higher the consequence, the less reasonable it is to outsource judgment to a probability score.
Vendors do not all use the same architecture, and proprietary systems do not expose every feature. These concepts explain the kind of evidence a text classifier may use, not a universal recipe.
A model can estimate how expected each next word is under a language model. Text with unusually predictable choices may receive a higher machine-likeness score, but polished human prose can also be predictable.
Some systems examine differences in sentence length, rhythm, vocabulary, and structure. This is sometimes described as burstiness. A genre, assignment rubric, editor, or second-language writer can naturally produce a consistent pattern.
A classifier can learn combinations of wording, syntax, punctuation, and document characteristics from labelled examples. The exact features are usually proprietary, and performance depends on whether the new text resembles the training data.
Length, language, formatting, code, quotations, lists, and mixed human-machine passages can affect the result. A paragraph-level score is not automatically transferable to an entire document.
Metadata, version history, signed records, or a watermark can provide information about origin when available. Provenance is different from guessing authorship from the final words alone.
A language model can assign probabilities to possible next tokens. Detectors built around this idea may look at how surprising or predictable the observed choices are under a model. Human writing that uses common phrasing can look predictable. AI writing that has been prompted for unusual style can look less predictable. The signal is useful for research and screening, but it is not unique to machine authorship.
Some detectors combine predictability with variation across sentences. A piece with very similar sentence lengths and repeated transitions may look different from a piece with sharp changes in rhythm. But genre creates these patterns too. A lab report, legal clause, job application, exam response, product description, and social post each have different conventions. Asking one classifier to make a universal authorship decision from all of them is a difficult generalization problem.
Modern detectors can use learned classifiers rather than exposing one simple rule. During development, the provider labels examples, trains a model to separate categories, evaluates it on a test set, and chooses thresholds for its product. The result depends on the examples, labels, language, model family, editing level, and distribution of the test set. A detector can perform well on one benchmark and poorly on the text a school, employer, or publisher actually receives.
The detector may also examine features that are not about authorship at all. Length, formatting, language, copied quotations, code, headings, lists, and the presence of mixed passages can change the input. A short answer contains less evidence than a long document. A translation may preserve ideas while changing the linguistic surface. A human editor can smooth writing until it resembles a different distribution. These are reasons to inspect the conditions before interpreting a result.
A detector can report accuracy, but the number is meaningful only with the test population, base rate, threshold, language, genre, model, and cost of each kind of error. In a disciplinary or employment setting, false positives and false negatives do not have equal consequences, so a vendor's headline score cannot answer the policy question by itself.
Human writing can be classified as machine-generated. A false positive is not a minor technical detail when a score affects a grade, job, reputation, or access.
Generated or transformed text can pass as human. A low score does not certify that a person wrote every word or complied with a policy.
A detector trained or tuned on one model, language, genre, or time period may not generalize to another. New models and editing workflows change the problem.
A document can contain a human outline, machine draft, human revision, copied quotation, translation, and another tool's edits. A single percentage hides that mixture.
Writing style correlates with language background, education, genre, and dialect. A system that mistakes standardized or less varied English for machine writing can create unequal error.
A percentage is not the probability that a named person used AI. It may be a model score under the vendor's test conditions, with a threshold chosen for a particular error tradeoff.
OpenAI's own 2023 classifier announcement said the classifier correctly identified only 26% of AI-written text in a challenge set and incorrectly labelled human text 9% of the time; OpenAI later removed it because of its low accuracy. That does not prove every current vendor has those exact rates. It does show why a confident interface should not substitute for validation in the context where a person may be punished.
Stanford researchers tested seven detectors on essays written by US-born eighth graders and TOEFL essays written by non-native English speakers. Their study reported that detectors classified 61.22% of the TOEFL essays as AI-generated, while 97% were flagged by at least one detector. The mechanism matters: features associated with less varied or more predictable English can be mistaken for machine writing. The exact numbers belong to that study and test setup, but the warning is broader than one product.
A fair policy must therefore ask who bears the error. If a detector performs differently across language backgrounds, dialects, disability-related writing patterns, or genres, an institution cannot hide that disparity behind a single threshold. The remedy is not to lower or raise the score until the desired outcome appears. It is to use better process evidence, test the tool on the affected population, give people an appeal route, and avoid treating style as character.
The same caution applies to human judgment. People are not reliable detectors either. Stanford research on human heuristics found that participants struggled to distinguish AI-generated self-presentations and relied on intuitive cues that could be manipulated. “It sounds like AI” is not a defensible evidence standard.
Use a detector, if the institution approves it, as a low-confidence prompt to inspect the student's process. Combine it with drafts, revision history, source use, and a conversation. Never make a disciplinary finding from a score alone.
Do not treat a resume or cover letter score as a proxy for honesty, ability, or suitability. Assess the candidate's work, ask consistent questions, and remember that spelling and grammar tools can change a document's signals.
A detector may help prioritize a quality or fact-checking review, but it cannot decide whether a writer complied with an outlet's disclosure policy. Check sources, reporting notes, interviews, drafts, and the agreed AI policy.
Use detection as a poor substitute for originality and usefulness. Review whether the content has evidence, firsthand detail, accurate claims, a clear audience, and an accountable editor. A low detector score does not make thin content valuable.
Use a fair process with notice, consistent evidence standards, confidentiality, and an opportunity to respond. A detector's probability is not an employment fact, and the consequences of a false accusation can be substantial.
For US college use, our AI detection guide for colleges goes deeper into process evidence, FERPA considerations, and student appeals. Our AI detectors overview is useful for comparing products, but a product comparison should not be confused with proof that a score is valid for your case.
“An AI detector may be used as one preliminary signal where appropriate, but its output will not be treated as proof of AI use or as the sole basis for a consequential decision. The organization will consider process evidence, apply the same standard consistently, protect personal information, and provide a meaningful opportunity to respond.”
Most AI writing detectors use a classifier trained on examples of human and machine-generated text. They may examine statistical signals such as how predictable the next words are, variation in sentence patterns, vocabulary, and other learned features, then return a probability or classification. The output is an estimate, not a direct observation of who typed the text.
No. A detector score by itself cannot prove authorship or intent. It can be a screening signal that prompts a fair conversation or a request for process evidence, but it can also miss AI text and falsely flag human text. High-stakes decisions should use corroborating evidence and an appeal path.
Human writing can resemble the patterns in a detector's training data. Predictable language, formal or standardized prose, short samples, editing tools, English-language learner writing, and some dialect or genre features can affect the score. Stanford research found serious false-positive concerns for non-native English writing.
Text transformation can change the statistical signals a detector uses, and detectors can also fail on lightly edited or newer model output. This is one reason a score should not be treated as proof. Detection systems and evasion methods change, so a result should be interpreted with the document's context and process evidence.
Use process evidence appropriate to the context: version history, drafts, notes, research sources, a short oral explanation, an in-class or supervised sample, code commits, an editorial trail, or a conversation about the argument and decisions. The evidence should be collected consistently and allow the person to respond.