← All posts

False Positive Rates in AI Content Detection Tools

Detectors flag predictable writing regardless of who wrote it, making false positives inevitable.

October 7, 2026

Cover illustration for “False Positive Rates in AI Content Detection Tools”

False positive rates in AI content detection tools are not a bug, and better calibration will not fix them. They are the predictable result of measuring sentence-level statistics that never had a reliable connection to who, or what, wrote the sentence.

What statistical signals AI detectors measure

AI content detectors do not read for authorship. They read for pattern, and they infer authorship from the pattern they find. The dominant method rests on two measures: perplexity, which tracks how surprising a sequence of words looks to a reference language model, and burstiness, which tracks how much that surprise varies from one sentence to the next across a passage. Human writing tends to swing between plain stretches and sudden, odd word choices. Machine output tends to stay smoother and more even, so a detector treats low perplexity and low burstiness as warning signs of machine origin.

A second method, watermarking, builds the signal in at the moment of generation. The KGW approach splits a model's vocabulary into a "green list" and a "red list" before a single word is generated, then nudges the model's internal scoring toward green-list words. A watermarked passage ends up using far more green-list words than chance would predict, and detection becomes a simple statistical test: count the green words, compare that count to what random chance would produce, and express the gap as a z-score, the number of standard deviations the actual count sits from the expected one. A z-score above roughly 4 triggers a flag.

Both approaches share one weakness. Low perplexity, high green-token density, and low burstiness all describe how predictable a piece of writing is. They say nothing, on their own, about who or what produced it. A detector built to catch predictable text will catch predictable text, regardless of its source.

Predictable Register and the Detection Threshold

Because the underlying measurement is predictability rather than origin, any human writer whose style leans on common words and steady sentence rhythm can produce a score that lands above the flagging line without a machine touching the draft. Structured academic prose, formulaic business writing, and the English written by non-native speakers all tend to share the same statistical fingerprint a detector is tuned to treat as suspicious: a narrow vocabulary, low perplexity, and sentence lengths that don't vary much.

Jonathan Karr and colleagues at the University of Notre Dame tested this directly. Their study, presented at AILS '26, ran unmodified peer-reviewed abstracts published between 2023 and 2025 through commercial detectors and found flagging rates of 9 to 15 percent on writing that no model touched. Non-STEM fields were flagged at far higher rates than STEM fields, and the researchers traced the elevated scores to long-token density and Academic Word List density, the kind of vocabulary that formal academic writing favors regardless of who wrote it. Authorship intent had nothing to do with the score.

A natural objection follows: vendors report very low false positive rates on their own tests, so why should a handful of academic abstracts overturn that record? Vendor benchmarks run on short, uniform, pre-labeled text built for clean testing conditions, while detectors are actually pointed at long, varied, real-world prose once deployed. A tool tested against the kind of text it was calibrated on will look accurate. The same tool, aimed at the kind of text people actually write for a living, behaves differently. The false positive problem here is built into what the measurement counts as evidence in the first place, not a tuning failure to be fixed with a better threshold.

Hybrid Human-AI Writing and the Detection Danger Zone

The writing most at risk of a wrongful flag is the hybrid draft, where a person writes the ideas and uses a language model to smooth the language, which is close to how most professional content gets produced today.

Karr and colleagues measured this directly. Their AILS '26 study found that "light refine" edits, where an author keeps their own analysis and uses a model only to polish phrasing, were flagged at rates of 38 to 80 percent, depending on which detector and which domain. A research team at the University of Maryland reported a related pattern: a single light editing pass through GPT-4o produced detection rates that swung widely from one detector to another. A writer doing nothing more than tightening a few sentences with an AI tool can see their own work flagged at rates that dwarf any vendor's advertised error rate.

The verdict also depends on which model did the polishing, not on how much the model actually changed. Text refined by older or smaller models gets flagged more often than text refined by newer systems. A writer's sanction risk, in other words, hinges partly on which tool happened to be open on their laptop that day.

The sharpest version of this problem is the one Karr's team documented directly: an honest writer who discloses a light AI-assisted edit faces a higher chance of being flagged than someone who runs the same text through a tool built purely to evade detection. After that kind of humanization pass, fewer than 4 percent of rewrites that had originally been labeled AI-generated still triggered a flag. The system built to catch AI use ends up punishing disclosed, limited assistance harder than it punishes concealment.

Domain and writing style as drivers of false flags

False positive risk does not spread evenly across writers. It concentrates on whoever's natural style happens to resemble the pattern the detector was built to catch.

Non-native English speakers carry much of that weight. Detectors calibrate on perplexity and burstiness patterns that non-native academic prose regularly produces on its own: a narrower working vocabulary, dependable connective phrases, grammar that is correct but doesn't vary much from sentence to sentence. None of that indicates machine authorship. All of it looks, statistically, like the thing the detector is hunting for.

Domain produces a parallel effect. The same AILS '26 research found non-STEM academic fields flagged at rates far above STEM fields, and the researchers found this gap statistically significant. The clear, structured prose that humanities writing favors raises a detector's score independent of who actually wrote it. Even the technical setup of the detector adds its own instability: the same passage can produce very different scores depending on how many tokens the tool analyzes in a single pass, a detail most detectors don't disclose to the people relying on their output, and no amount of threshold adjustment resolves it.

Put together, two writers can submit equally original work, in different fields or different registers, and walk away with opposite verdicts from the identical tool. The score tracks style and domain far more reliably than it tracks whether AI touched the page.

The structural gap between vendor benchmarks and real-world performance

The distance between a vendor's advertised false positive rate and what independent researchers actually find is what happens when a tool is scored on one kind of text and then sold for use on an entirely different kind.

Vendor benchmarks run on text built for testing: short, consistent in style, and labeled with certainty as human or machine before the test even starts. Real corpora, the kind an enterprise content team or a university actually produces, look nothing like that. They run long, cross multiple domains, shift in register from paragraph to paragraph, and often blend human and AI contributions inside the same document. The AILS '26 study found that commercial detectors, run against real published abstracts using proxy ground-truth labels, produced false positive rates far above the figures vendors report, and that those rates shifted significantly by domain and by publication year in ways no benchmark would ever surface.

Detectors also disagree with each other on the same piece of writing. The same passage can receive wildly different AI-likelihood scores depending on which tool reads it, a sign that the number reflects how a given detector was built more than any fixed property of the text itself. A research team at the University of Chicago Booth reported that detection accuracy swings widely depending on the tool, the length of the text, which model generated the original draft, and where the threshold is set. No single standard holds across those variables. A claim of high accuracy only describes performance on the one benchmark it was measured against. Tuning the threshold inside that benchmark will never fix a gap that exists because the benchmark and the real world aren't drawn from the same population of text.

Institutional Reliance on Detector Scores: Karr et al. and Notre Dame

An institution that treats a detector's number as proof of AI authorship is stating a fact the underlying science cannot support. Karr and colleagues make this point without hedging: detector scores should not stand alone as misconduct evidence, because the false positive rate on guideline-compliant AI assistance runs too high, and the false negative rate on humanizer-assisted evasion runs too low, for the number to carry the evidentiary weight institutions ask of it.

A concrete case shows what that looks like in practice. In Matter of Newby v. Adelphi University, decided by the New York Supreme Court in Nassau County in January 2026, a student was marked for academic misconduct after Turnitin returned a "fully AI-generated" score. The student disputed the finding and produced results from other detectors that classified the same text differently. The court annulled the academic-integrity finding and ordered the record expunged. The ruling did not declare the detection model scientifically wrong; it found the evidentiary and appeal process the university relied on insufficient to support the sanction. The distinction matters for any institution weighing how to use these tools: the court didn't need to settle the science to find the process unsound. An institution doesn't need to wait for a scientific consensus to recognize the same procedural exposure in its own policies.

Scale turns that exposure into something harder to absorb. An institution running detector checks across a large volume of assessments every year accumulates a systematic stream of wrongful findings, each one requiring individual investigation to unwind.

Detection reliability in multi-model pipelines

Enterprise content work rarely passes through a single model anymore. A draft might move through one model for tone, another for structure, another for fact-checking, and each handoff changes the underlying word choices in ways a single-model detector was never built to track. That kind of pipeline doesn't just make detection harder. It dissolves the statistical signature that KGW-family detectors are built to find, so the detector's output stops functioning as a meaningful signal of where the text came from.

Watermarking and perplexity-based detection both assume one stable source producing one consistent pattern of word choices. Research on the Bias Inversion Rewriting Attack, or BIRA, shows how easily that assumption breaks. By suppressing green-list words while keeping the meaning of a passage intact, researchers dropped the z-score of KGW-watermarked text generated by Llama-3.1 from 6.03 to 0.83, below the usual detection threshold, without any noticeable distortion of what the text actually said. A pipeline that routes a draft through several different models produces a similar dispersal as a byproduct of how it works, with no adversarial intent required.

That leaves enterprise teams exposed on two sides at once. A passage can get flagged because one model's word choices happen to look suspicious on their own, or it can clear detection because the mix of models accidentally scatters the exact signal a genuinely AI-written passage would otherwise carry. There is no reliable way to know, from the score alone, which of the two has happened.

Lettershred's architecture addresses that problem at its source rather than treating it as something to patch after a draft is finished. It routes every line of a piece across many model kernels at once, with a panel of judges voting line by line on which output to keep, so the resulting mix of word choices is genuinely spread across multiple providers. The z-score that results falls below the detection threshold because no single model's fingerprint dominates the final text, not because anyone edited the surface of the writing after the fact.

Treating a single detector verdict as a signal, not a finding

None of this argues for abandoning detectors. It argues for treating a detector's output as one data point among several, never as the sole basis for a compliance call, a policy decision, or a misconduct finding. The false positive rates documented in hybrid human-AI writing, the bias toward certain domains and writing registers, the disagreement between tools reading the same passage, and the instability introduced by something as basic as window size all point the same direction: a detector's score reflects how that detector was designed as much as it reflects any real property of the text in front of it.

Running a passage through several detectors and looking for agreement narrows the problem without solving it. Tools that disagree on the same text are proof that no shared, objective standard exists behind any of them, and a consensus built from several flawed instruments is still a flawed verdict, just one with more confidence attached to it than it deserves.

Because detectors measure statistical properties rather than authorship itself, tools built on architectural diversity, Lettershred among them, blend output from multiple language models to scatter the single-model fingerprint before detection ever begins, addressing the root cause. The paradox Karr and colleagues documented, that honest writers doing light AI-assisted edits get flagged more often than writers using tools built purely to dodge detection, points to why timing matters more than cleanup: an approach that removes the detectable single-model signature at the moment of generation lets a content team keep its legitimate workflow intact while avoiding that same risk.

Lettershred's own pipeline illustrates the mechanism. A brief goes in, a dozen model kernels draft in parallel, a panel of judges picks the strongest line from among them, and a final pass unifies the voice across the piece. Measured against real KGW detectors, the resulting z-scores have dropped from 15.46 to -0.32, with most of the underlying tokens rewritten through that cross-provider mix, putting the output below the detection threshold by the way it was built.

The underlying measurement problem persists even when one pipeline handles it well, so the durable response to a structurally unreliable test is a production process built so it never has to pass that test.

AI Watermark Detection

Sources