Artificial text detection and artificial text generation are two faces of a single machine. A language model is trained to predict the next word, and a detector tries to read the statistical residue that prediction leaves behind. To understand why a perfectly honest essay can be branded "AI," one has to follow the whole chain: how these models are built, what they read while learning, the habits that training leaves in their prose, and the blunt instruments we use to detect those habits afterward.
A modern language model — GPT, Claude, Gemini, Llama — is a neural network built on the Transformer architecture introduced in 2017 in a paper memorably titled "Attention Is All You Need."[1] Its defining mechanism, self-attention, lets every word in a passage weigh how strongly every other word should influence what comes next. That single idea is what allowed models to scale from sentence-completion toys into systems that draft essays.
Training unfolds in two great phases. In pre-training, the model is shown an enormous corpus — hundreds of billions of words drawn from the open web, digitized books, and reference sites — and given exactly one objective: predict the next token. A token is a word or a fragment of one. Across trillions of guesses, corrected each time against the real next token, the network slowly encodes grammar, facts, and rhythm, not as rules but as probabilities. Nothing tells it what an essay "should" sound like; it simply learns what text statistically tends to do.
The second phase, fine-tuning with human feedback (often called RLHF, Reinforcement Learning from Human Feedback), is where the recognizable personality appears. Human raters score candidate responses, rewarding answers that feel helpful, balanced, and polished.[2] The model learns that a hedged, smooth, symmetrically organized paragraph earns a high score. Many of the "AI tells" people notice are not accidents; they are the taste of thousands of raters, distilled and amplified.
Here lies one of the most counter-intuitive facts in this whole subject. Encyclopedic and reference writing — Wikipedia above all — forms a large, clean, heavily weighted slice of pre-training data.[3] The calm, neutral, "explainer" voice that models default to was learned substantially from exactly that material.
The consequence is circular and a little absurd: Wikipedia-style prose is often scored as "AI-like" because the AI was trained to write like Wikipedia. A detector trained to notice the model's neutral register will, naturally, flag the very source that taught the model that register. The same trap catches any clean, encyclopedic, well-organized human writing — which is to say, it catches good students. Other heavily weighted sources leave their own marks: academic papers contribute hedging and formal transitions; marketing copy donates buzzwords like leverage, robust, and seamless; technical documentation supplies the tidy, list-like scaffolding that fine-tuning later rewards.
When the model writes, it turns the running text into a probability distribution across its entire vocabulary and samples one token, then repeats the whole computation for the next. Three controls shape the result. Temperature governs adventurousness: low values make the model choose the single likeliest word, producing safe and predictable text; high values flatten the odds and invite surprise. Top-k and top-p restrict sampling to the most probable handful of candidates.
Because the safe, default setting is low-temperature and on-distribution, the output naturally gravitates toward the most statistically probable phrasing available. This is the deep reason AI prose feels smooth, and it is precisely what a detector measures as low perplexity: text that never surprises a model is text that walked the path of maximum probability.
Researchers at Stanford and the University of Tübingen found that vocabulary choice alone predicts AI authorship with better than seventy percent accuracy, before any structural analysis at all.[4] A separate study published in the Proceedings of the National Academy of Sciences isolated structural habits, the strongest of which is a small grammatical quirk most readers never consciously notice.[5]
The clearest tell is vocabulary. A cluster of words appears in machine text at many times their human rate: delve, which appears roughly forty-eight times more often, along with tapestry, multifaceted, leverage, robust, utilize, navigate, underscore, and seamless. A sentence such as "Let us delve into the multifaceted landscape of this nuanced issue." trips three of them at once.
Close behind comes the family of hedges and fillers the model uses to sound careful. Consider the difference between two openings of the same idea.
The first hedges, abstracts, and reaches for stock phrases; the second commits to a claim and supplies a concrete image. Detectors cannot read meaning, but they can count the hedges. A third tell is structural symmetry: triadic lists such as "social, economic, and political," paragraphs of suspiciously even length, signpost introductions, and formulaic closers. The fourth, and the strongest structural signal in the research, is participial padding — the trailing "-ing" clause that adds grammar but no information. In "The policy reduced emissions, highlighting its importance in the broader effort," the words after the comma are pure decoration, and models produce them at several times the human rate. Finally there is rhythm: human writers mix a four-word sentence against a thirty-word one, while models hold an even, metronomic cadence. This is burstiness, and its absence is one of the most reliable signals there is.
Commercial and academic detectors fall into three families, and all of them read statistical residue rather than verifying who actually held the pen. The first is perplexity-based, the approach that made the original GPTZero famous.[6] The detector runs the candidate text through its own reference model and measures how surprised that model is at each word. Take the sentence "The mitochondria is the powerhouse of the cell." Every word is the overwhelmingly likely continuation of the one before it, so the model registers almost no surprise — low perplexity — and the sentence reads as machine-smooth, even though a million students have written it by hand. That false positive is the method's signature weakness.
The second family is burstiness-based. It ignores individual words and measures the variation in sentence length and structure across the passage. Given five consecutive sentences of eleven, twelve, ten, eleven, and twelve words, the detector sees a flat line and leans toward "AI." Given sentences of four, twenty-six, seven, and nineteen words, it sees the jagged profile of a human mind changing pace. The weakness is equally clear: a disciplined writer trained to keep sentences uniform, or a model deliberately told to vary them, defeats the measure.
The third and now dominant family is classifier-based, used by tools such as Originality.ai.[7] Here a machine-learning model is trained on enormous labelled collections of human and AI text and learns the vocabulary and structural distributions directly. It is the most accurate family on clean, unedited AI output and the most opaque: it returns a confidence number with little explanation, and it inherits every bias in its training set. A fourth approach, watermarking, is provider-side rather than detector-side — the model nudges its own token choices toward a secret statistical pattern that a matching verifier can later recover.[8] Ordinary detectors cannot see it.
The root problem is unavoidable. Because models are trained on human writing, they are built to imitate human patterns, so the better and cleaner a human writes, the more the detector's signals fire. For the vast middle of real writing — competent, formal, unremarkable schoolwork — the signals simply collapse. A plain student essay and a machine essay land in the same ambiguous band, because neither carries the vivid specifics or jagged rhythm that would distinguish them. Add the ease of evasion (a single pass through a paraphraser, a translation, a raised temperature) and the fact that short passages never carry enough signal, and the verdict becomes a coin-flip dressed up as a percentage. A 2025 review of academic detectors found accuracy varied widely and no tool reached reliability, while false positives carried career-altering stakes.[9]
No group is harmed more by these tools than writers for whom English is a second language. A Stanford study fed TOEFL essays written by non-native speakers into seven leading detectors; more than half of the genuine human essays were flagged as AI, while essays by native speakers passed almost untouched.[10] The mechanism is cruel in its simplicity. Second-language writers tend to deploy a narrower, more common vocabulary and more regular sentence structures — exactly the low-perplexity, low-burstiness profile that detectors read as machine-made. The very features that mark careful, learned, non-native English are the features the detector punishes.
There is a slower, stranger problem on the horizon. As people read ever more AI-generated text — in search results, emails, articles, and homework help — their own writing drifts toward the machine's register. Linguists call this linguistic homogenization or language flattening: a measurable narrowing of vocabulary and syntactic variety across a population. Early studies of spoken and written language after the mass adoption of chatbots already show usage spikes in signature words like "delve."[11]
The feedback loop is self-defeating for detection. If models learned their voice from humans, and humans now learn their voice partly from models, the two distributions converge. A detector that works by spotting the gap between human and machine writing has less and less gap to find. The honest conclusion is that these signals never operate alone. Combined with the ESL penalty, the convergence of human and machine style, short samples, topic effects, and the simple fact that good formal writing is supposed to be clean, they compound into a high rate of confident, wrong answers.
It would be a mistake to read all of this as an argument against the technology. For a second-language learner, a language model is among the best writing tutors ever built: patient, instant, available at midnight, willing to explain why a preposition is wrong for the fortieth time. Used well, it can lift a student's clarity and confidence enormously.
But the same tutoring quietly pushes the student's prose toward the model's own register. A learner who drafts in their own voice and then asks an AI to "make this sound more academic" will receive their ideas back wrapped in leverage, multifaceted, and "it is important to note." They have not cheated; they have learned from the tool exactly as intended. Yet their essay now carries the fingerprints a detector hunts for. This is the deepest unfairness in the system: the better a tool teaches a struggling writer to sound "academic," the more likely that writer is to be falsely accused, because the academic register the tool teaches is the register the tool itself writes in. Detection, in the end, cannot be separated from pedagogy, equity, and the slow drift of language itself — which is exactly why a score should inform a conversation, never end one.