← All posts

Z-Score Thresholds in AI Watermark Detection

Thresholding decisions determine what counts as AI-generated, not just detection accuracy.

September 28, 2026

Cover illustration for “Z-Score Thresholds in AI Watermark Detection”

The z-score threshold is the single number that decides whether a piece of text gets flagged as AI-generated or waved through, and understanding how that number gets built is the difference between reasoning about watermark risk and just guessing at it. That's the entire trick. Watermarking doesn't rewrite what a model wants to say, it just tilts the odds on which words it reaches for. Z-Score Thresholds in AI Watermark Detection.

Detection reverses the process. A verifier with the secret key can reconstruct, token by token, which words should have counted as green, then count how many actually showed up in the candidate text. The formula for that count is z = (|x|_G − γT) / √(Tγ(1−γ)), where the numerator captures how far the observed green-token count sits above what pure chance would produce, and the denominator is the standard deviation you'd expect under the null hypothesis that no watermark exists at all. Divide one by the other and you get a z-score: a single number expressing how many standard deviations away from "ordinary human text" a document sits.

What makes this elegant, and worth respecting as a piece of engineering, is how little the detector needs to know. No model weights, no access to the probability distribution the generator used, nothing proprietary. Just the token identities in the text, the secret key needed to rebuild the green list, and the length of the text. Every term introduced here, green tokens, red tokens, γ, δ, z-score, threshold, carries through everything that follows. The KGW scheme originates with Kirchenbauer et al., who partitioned the vocabulary into green and red subsets, giving green tokens a positive logit bias (δ) at generation time.

Why z₀ = 4 is the conventional threshold and what it guarantees

Four. (2023a) for KGW and related schemes.

A companion metric, TPR@5% FPR, is the true-positive rate a scheme achieves once its threshold is calibrated so that exactly 5% of genuinely unwatermarked text gets falsely flagged. This has become the standard yardstick for comparing removal attacks and detection schemes against each other, because it holds the false-positive side constant and asks only how much true signal survives.

Within this framework, three tiers describe what a z-score actually communicates. Above 10, you're looking at high-confidence detection, the kind of score clean, unattacked watermarked output typically produces. Between the threshold and 10, detection is still positive but the signal is weaker, more fragile, more sensitive to whatever happens next to the text. Below the threshold, the detector clears the text, full stop, regardless of where it actually came from. In practice, average z-scores for watermarked text run well past 10, while non-watermarked text clusters near zero, which is exactly the wide separation the scheme was designed to produce.

None of this is physics. z₀ = 4 is a policy decision, and any operator running a detector can move it up or down depending on how much false-positive risk they're willing to swallow versus how many real watermarks they're willing to miss. A concrete example from signature-filtering research makes the stakes tangible: one unwatermarked passage scored z = −0.14, comfortably in "human" territory, while a watermarked passage scored just z = 1.63 without any additional processing, still below the z₀ = 4 line, and would have cleared the detector entirely. Applying a filtering step to that same watermarked passage causes its score to jump to 6.67, decisively above threshold. Same threshold, same underlying text, opposite verdicts, depending entirely on how much of the signal survives to be measured. The threshold z₀ = 4 controls the false positive rate (the probability of flagging genuinely human-written text as AI-generated), and setting the threshold here keeps that rate very low.

What moves the z-score: token length, temperature, γ, and δ as the four control variables

A z-score is not a fixed property of a document, it's an output of several interacting variables, and the first is simply length. The denominator of the z formula grows with the square root of text length, while the numerator, the excess of green tokens, grows linearly whenever a genuine watermark signal is present. Longer documents therefore accumulate evidence faster than they accumulate noise. Short outputs are inherently harder to detect with confidence, and long ones tend to produce decisive scores.

γ, the green fraction, sets the baseline. Widening the green list makes it easier for a model to land on green tokens simply by chance, which compresses the gap between watermarked and unwatermarked behavior even as it makes the watermark easier to embed in the first place. There's a tension built into the parameter itself.

δ, the logit bias, is the most direct lever. Pushing δ higher strengthens green-token preference, raises z-scores, and makes detection more reliable, but this comes at a real cost to output quality. Under realistic hyperparameter settings, watermarking has already been shown to cause drops of 10 to 20% in classification tasks and 5 to 15% in long-form generation quality, with every point of δ spent buying detection confidence a point spent degrading the very language model output the watermark is supposed to protect. That's not a rounding error. Every point of δ spent buying detection confidence is a point spent degrading the very language model output the watermark is supposed to protect.

Temperature completes the set, and it cuts in a counterintuitive direction. Low temperature concentrates the model's choices on its highest-probability tokens, which weakens the accumulated green-token signal and pulls the z-score back toward zero, making the watermark harder to detect. Crank temperature up, and the model explores more randomly, which strengthens the watermark's statistical footprint, but at the risk of degrading coherence and quality. Detection and fluency pull in opposite directions here, and there's no setting that maximizes both.

Entropy-Weighted Detection, or EWD, doesn't touch generation at all but reweights individual tokens by their entropy at detection time, aiming squarely at the tension between detectability and language quality. The same green-red generation scheme produces the underlying signal, and smarter scoring is layered on top of it to read that signal. Across all four variables, the z-score a document receives isn't baked in at the moment of generation; it's a function of configuration choices made before generation, the length of what gets produced, and the scoring method applied after the fact.

The dilution problem: how mixing watermarked and non-watermarked text collapses the z-score

Standard KGW detection collapses an entire document into a single global z-score. That design choice works fine when the whole document came from one watermarked model. It breaks down the moment only part of it did.

Picture a report where a single AI-drafted paragraph sits inside pages of human-written material. The green-token signal from that one paragraph gets averaged across every surrounding human token, which drags the numerator down without shrinking the denominator by anything close to the same proportion.

Researchers behind WaterSeeker, a detection method from Tsinghua University and collaborators published at NAACL Findings 2025, state that existing full-text detection methods "fail in real-world scenarios where LLMs generate only brief segments within longer documents, due to dilution effects". Their fix is a two-stage approach, first locating suspicious regions through anomaly extraction, then running local traversal and full detection only on the candidate window rather than averaging across the whole document. A separate approach, the WinMax algorithm from Kirchenbauer and colleagues, tackles the same problem by brute-force testing every possible window size and keeping whichever score is highest, though at a real computational cost.

None of this is an edge case dreamed up for a paper. Any real document that combines AI-drafted sections with human editing, legal boilerplate, or quoted material faces this exact scenario. And the practical consequence cuts both ways for any team relying on these tools: dilution can cause genuine AI content to escape detection, and it complicates the reliability of any positive or negative result a detector returns.

How paraphrasing and targeted rewriting push z-scores below the detection threshold

Attacks against watermark detection split into two categories. Query-based attacks probe a model directly, issuing crafted prompts to infer the green-token list, effective but demanding real access to the target system. Query-free attacks operate purely on the finished text, in a black-box setting, usually through LLM-based rewriting, and this is the category that matters most in practice, because it requires nothing more than a second language model and the original output.

The watermark's own design creates the opening these attacks exploit. Green-token bias has the most influence precisely where the underlying model is most uncertain, at high-entropy positions, so the watermark's signal concentrates exactly where a paraphraser has the most room to substitute one plausible word for another.

A method called BIRA, Bias-Inversion Rewriting Attack, presented at ICML 2026, exploits this directly. It applies a negative logit bias to a proxy suppression set identified through token surprisal, and critically, it never needs to know the actual secret green list the original watermark used. Reducing the average conditional probability of sampling a green token by even a small margin causes detection probability to decay exponentially as a function of text length and δ squared. Small, consistent pressure compounds fast. The reported results back this up, with evasion rates exceeding 99% across a range of watermarking schemes, while preserving semantic fidelity noticeably better than earlier rewriting attacks managed.

The takeaway is blunt: a slight, systematic nudge away from the green-token distribution, sustained across a paragraph, is enough to drag a z-score below the z₀ = 4 line. An attacker doesn't need to reverse-engineer the key or reconstruct the exact green list, approximate pressure in the right direction is sufficient. Earlier query-free methods, which worked by masking and regenerating high-entropy tokens, exploited the same underlying weakness but did so less efficiently. BIRA's contribution is less a new vulnerability than a sharper, more principled way of hitting the one that was already there.

What robustness benchmarks show when detectors meet real attacks at scale

The most thorough stress test to date is WaterPark, published by Liang and colleagues at EMNLP Findings 2025, which ran 10 watermarking methods against 12 attack types, across three language models (OPT-1.3B, LLaMA3-7B, and Qwen2.5-14B) and five datasets. The scale of that study makes its findings hard to wave away.

SynthID, under nothing more aggressive than moderate paraphrasing, saw its true positive rate fall from 0.998 on clean text to 0.498, roughly half its clean-text performance. A single paraphrase pass through ChatGPT was enough to drop every method tested in the benchmark below 30% detection. KGW itself showed how unstable a fixed threshold can be across models: true positive rate of 0.858 on OPT collapsed to 0.334 on LLaMA3 under identical attack conditions.

A separate forensic evaluation, run by Tamim and colleagues in a preprint submitted to AIES 2026, pushed further: across 846 valid paraphrase runs spanning 15 prompts per method, every KGW and Unigram text that had initially been detected lost its watermark entirely after paraphrasing, a 100% conditional removal rate, with SynthID close behind at 98.3%. And that's before accounting for the baseline failure rate: pre-attack false-negative rates, meaning failures on clean, unaltered text with no adversarial action taken at all, ran to 70% for KGW, 83% for Unigram, and 80% for SynthID. The majority of watermarked output in that study failed detection before anyone tried to break it.

In the same evaluation, 5.4% of paraphrased human-written control text got flagged as AI-generated, producing an 18.6% overall paradox rate for SynthID, and 80% of SynthID's own pristine watermarked output landed in an uncertain middle zone rather than a clean detection. None of this proves watermarking is a wasted effort. It proves that a single-scheme detector built around one global z-score is structurally brittle, and any team treating the absence of a flag as proof of human authorship is standing on ground that gives way easily.

Detection improvements that work within the z-score framework rather than replacing it

The response from the research community hasn't been to abandon the z-score, it's been to feed it better inputs. Signature filtering, developed by Hong and colleagues at National Chengchi University in a paper posted to arXiv in June 2026, adds a detection-time module that strips out pre-computed "signature" tokens, tokens whose presence tends to make the z-test unreliable, before the test is ever run. In weak-signal and low-entropy conditions, where detection was previously landing somewhere between 8% and 31%, filtering pushed detection rates up to between 78% and 99%, while keeping false positives under control. The concrete figures cited earlier make the mechanism visible: filtering pushed a watermarked passage's score from 1.63 to 6.67 while nudging an unwatermarked passage only from −0.14 to 1.66, widening the separation between the two from 1.77 to 5.01, enough to flip the detector's verdict at the z₀ = 4 line. Stress testing under scrambled sentences and 25-50% token perturbation (dilution, deletion, substitution) showed the filtering approach held onto most of its clean-text detection gains even under that pressure.

A second approach, the Pattern Stability Score, from Mansouri and colleagues at George Mason University, takes a broader view. Rather than relying on a single global z-score, PSS combines global and local z-score features with run-length pattern statistics, autocorrelation signals, and stability measures tracked across paraphrase depth, improving detection AUC by 10 to 15 percentage points over plain z-score thresholding. A single universal classifier built this way held above 87.8% AUC even when the generating model, the paraphrasing tool, and the text domain all differed from what it was trained on, which is a meaningfully different claim than most narrow benchmark results make.

A third direction changes what gets partitioned rather than how it gets scored. Topic-Based Watermarking, presented at ACL 2026 Findings, clusters vocabulary into green lists by semantic topic instead of random assignment, holding perplexity close to unwatermarked baselines while improving robustness against paraphrasing and lexical substitution. The common thread across all three approaches: none of them throw out the z-score framework. They augment it with richer signal, and critically, every one of them operates on the detector side, requiring no changes to whatever generation pipeline is already running in production.

Why a structurally single-model pipeline is vulnerable even when no active attack occurs

Setting aside the attackers and the paraphrasers, a quieter problem remains. Every output from a single watermarked model carries the same statistical fingerprint, the same green-list bias applied consistently, sentence after sentence, because it's all generated under the same key. That consistency is precisely what makes detection possible in the first place, and precisely what makes it so exposed.

Repetition compounds this. The more content an organization publishes from one model, the more its aggregate statistical pattern reinforces itself across the body of work, raising the effective z-score any sampled document is likely to produce. Nobody has to attack anything for this to matter. It's a property of homogeneous generation itself.

Adding the quality cost on top makes the picture worse. The 10 to 20% average degradation in classification tasks under realistic watermarking hyperparameters means a single-model pipeline is producing text that is simultaneously easier to flag and measurably worse. Paraphrasing after the fact doesn't solve this, it's the mechanism attackers already use, and it leaves the underlying token-selection statistics largely intact unless the rewriting happens across genuinely different model providers rather than as a cosmetic pass. The z-score threshold, in this light, isn't mainly a risk imposed from outside by regulators or third-party detectors. It's a risk a homogeneous pipeline manufactures for itself, and no amount of manual editing changes the architecture that produced it.

How multi-model authorship disperses the green-token signal below the detection threshold by construction

The fix follows directly from the diagnosis. If a document is assembled from multiple models, each operating under a different provider's key and a different green-list partition, no single green-list pattern can dominate the aggregate text. The z-score formula assumes one consistent γ and one consistent key running through the whole sample, and cross-provider authorship breaks both assumptions at once, not as an afterthought but by the basic arithmetic of the test itself.

Lettershred's production architecture is built around exactly this principle. A brief enters the system, dozens of model kernels draft candidate lines in parallel, a panel of roughly a dozen judges votes line by line to select the strongest version, and a final enhancement pass unifies the surviving lines into one coherent voice. Diversity of authorship at the token level, one voice at the surface. That's the design, not an accident of process.

The measured result: output z-scores fell from 15.46, comfortably in high-confidence detection territory, to −0.32, below the null expectation for ordinary human writing, with 177 of 240 tokens in the sample rewritten through this cross-provider dispersal. A score of −0.32 isn't a near-miss on the threshold. It's on the other side of zero entirely, which means the text doesn't read as weakly watermarked, it reads as statistically indistinguishable from text no single model wrote. That's compliance built into the architecture, not achieved through a lucky paraphrase or a manual edit pass applied after the fact.

AI Watermark Detection

Sources