Green-List Token Bias and Why It Accumulates
The bias compounds across thousands of tokens.
October 2, 2026

Green-list token bias is not a dramatic intervention somewhere in a generated text; it is a mechanical nudge applied at every single token, and that nudge is why the signal grows the longer a model keeps writing.
How the KGW partition works at the moment a token is chosen
Before a language model produces its next word, the KGW scheme runs a quiet partitioning step that has nothing to do with meaning. The full vocabulary gets split into two groups, a green list and a red list, using a hash of the token that immediately precedes the one about to be generated. That hash seeds a pseudorandom process, so the split is not fixed in advance and does not correspond to any stable idea of "preferred" vocabulary. A word that lands on the green list after one context token can just as easily land on the red list after a different one. This context-dependency is precisely what makes the mechanism hard to notice by reading: there is no blacklist of banned words to spot, because the lists themselves are reshuffled at every step.
Once that split exists, the scheme adds a positive bias, denoted δ, to the logits of every green-listed token before the model converts those logits into probabilities and samples from them. The effect is not to force a particular word; it is to lean the odds toward the green half of the vocabulary. Wouters (2024) frames this as a trade-off: a stronger δ makes the watermark easier to detect later, but it also makes the text more likely to deviate from what the model would otherwise have produced, and that tension sits at the center of the scheme's design and cannot be engineered away.
The green list carries no semantic judgment. It does not know what a word means, only which side of a hash function it falls on, and that blindness has a real consequence when very few plausible continuations exist. Consider a sentence that calls for the name of a specific public figure, where almost any other word would be an obvious error. If the one correct answer happens to fall on the red list, the model is pushed toward a less natural, less accurate green alternative instead. The research brief gives a concrete example: asked to continue a sentence about the CEO of SpaceX, "Musk" may simply fall on the red list, forcing a degraded or evasive continuation rather than the correct name. Looked at on its own, this single substitution at a single token looks like a minor curiosity, almost too small to matter. What happens when that same small lean gets applied thousands of times in a row is the real question.
Why the bias accumulates rather than averaging out
A thumb permanently resting on one side of a scale does not average out no matter how many times the scale is used, and that is the right way to picture what the KGW bias does across a generated text. Because the hash function is deterministic, the same preceding token always produces the same green-red partition, so every sampling event performed under the same watermark key pushes probability mass in a consistent direction rather than a random one. STELA's authors identify this determinism as the exact property that KGW exploits to make its signal detectable in the first place. The bias is not a coin flip that sometimes favors green and sometimes favors red across a long document; it is one directional lean, reapplied at every single step.
Detection exploits that same consistency. A detector counts how many tokens in a suspect text fall on the green side, compares that count to what chance alone would predict, and expresses the gap as a z-score, with the text's length and the green list's expected proportion built into the normalization. Because the z-score grows with the number of tokens examined, a longer text carries more statistical power: the very same per-token bias δ that is nearly invisible in a single paragraph becomes overwhelming across a full article. Length is not a neutral variable here; it is the mechanism's natural ally. A short passage can sit comfortably below a detection threshold while the identical underlying bias, applied across several thousand words instead of a few hundred, pushes the z-score decisively above it.
The accumulation is also uneven in a way that matters a great deal downstream. It gathers most strongly wherever the model genuinely had more than one reasonable option and almost not at all wherever only one word was ever plausible. The KGW approach performs well on open-ended, expressive writing and performs poorly on constrained output like code generation or machine translation, where correct answers are narrow and the room for a green-favoring nudge to operate is correspondingly small, a limitation the research brief documents directly.
Entropy and the Signal's Shape
The unevenness just described has a name in information-theoretic terms: entropy, and it determines where the watermark's signal actually lives inside a text rather than how strong that signal is on average. High-entropy positions are the forks in the road, points where several candidate tokens are at roughly equal probability, synonyms, optional connectives, alternative phrasings that would all read naturally. At these forks, δ reliably tips the outcome toward whichever option happens to be green, even though a red-listed alternative would have served the sentence just as well. Multiplied across hundreds of forks in a long document, that tipping becomes a cumulative preference that is statistically unmistakable even though no individual substitution looks wrong.
Low-entropy positions behave in the opposite way. Where one token is overwhelmingly more probable than any alternative, two outcomes are possible, and neither resembles what happens at a high-entropy fork. If that dominant token happens to be green already, it gets chosen as it would have been anyway, adding to the green count without distorting anything. If that dominant token happens to be red, the model gets pushed toward a lower-probability green substitute, and this is precisely where factual accuracy and fluency come under strain, as the earlier SpaceX example demonstrates. STELA was built specifically to correct this asymmetry, adjusting watermark strength according to how predictable a position is syntactically rather than applying a flat δ regardless of context.
A text's detectability and a text's damage come from two different populations of tokens. A long document's z-score is carried mostly by its high-entropy positions, while whatever quality cost the watermark imposes concentrates instead at its low-entropy positions, and these are genuinely separate problems that do not track each other. That separation explains an otherwise confusing fact: a text can show a low increase in perplexity while still carrying a strong, easily detected watermark signal, or it can show a high perplexity increase while carrying only a weak one. Perplexity auditing and z-score detection are measuring different things, and a team relying on only one of them is seeing only half the picture. This same unevenness also explains why the signal tends to survive casual editing, since most edits land on high-entropy, stylistically flexible passages where the watermark is thickest, while it stays exposed to anyone who specifically targets those same high-entropy forks for suppression.
How the bias distorts text quality in ways that are invisible without instrumentation
The quality cost of green-list bias does not announce itself in any single sentence, which makes it a consistent pull away from whatever word the model would have chosen on its own and toward whichever green-listed alternative clears the δ threshold at that position. Across a long piece of writing, this produces a text that reads as faintly "off" without any one passage being identifiably wrong, because the aggregate pattern of word choices was never the pattern the model would have produced freely. The research brief describes the effect precisely: the KGW method is "continually shifting the generated answer from its original distribution through rigid and rough vocabulary splitting".
For an enterprise content team trying to hold a consistent brand voice, this distortion has a concrete operational cost. If a watermarked model systematically favors one synonym over an equally valid alternative, or one sentence structure over its equivalent, sentence after sentence, draft after draft, the cumulative result is a house style the team never actually chose and would not recognize as deliberate. Reading the output will not reveal this drift, because each individual substitution is locally plausible on its own terms; only perplexity auditing surfaces the pattern across many substitutions. The cost is worst at low-entropy positions, where the forced green alternative is not a harmless stylistic variant but a genuinely lower-probability, sometimes factually weaker or semantically imprecise substitute. The research brief names code generation and machine translation specifically as domains where this failure mode occurs acutely.
SynthID-Text, published in Nature by Google DeepMind in October 2024, was built in direct response to this class of problem, using a sampling-based approach rather than logit biasing, which makes it non-distortionary by design rather than merely better-tuned. That design choice sharpens rather than removes the underlying tension. The SynGuard robustness assessment documents that SynthID-Text's detection accuracy drops sharply under moderate paraphrasing or back-translation, so the gain in preserved quality comes paired with a cost in robustness. Wouters's framing of the trade-off applies again here in full force: no scheme simultaneously maximizes quality, detectability, and robustness at once. The quality distortion inherent to logit-biasing approaches like KGW is not a flaw that a future patch will quietly fix; it is a structural feature of how the approach achieves detectability in the first place.
Why the bias becomes more legible when the same pipeline runs many drafts
Everything true of a single long document becomes true, in a different and more consequential way, of an enterprise pipeline that generates many documents under the same watermark key. What grows with the length of one text also grows with the volume of texts produced by the same model under the same conditions, and at that scale the accumulated bias is no longer a property of individual documents but a property of the model's output distribution itself. Water-Probe, presented at ICLR 2025, demonstrated that current watermarked language models expose consistent biases under a shared watermark key, and that almost all mainstream watermarking algorithms can be identified through carefully designed prompts: the underlying fingerprint is stable enough to be learned purely from black-box sampling, without any access to the model's internals. Even when no single watermarked text looks perceptibly different from an unwatermarked one, the distribution across many watermarked texts reveals whether the model producing them is watermarked at all.
A content operation that routes many briefs through a single model is, functionally, repeatedly sampling under the same key, and each additional draft makes the aggregate fingerprint more legible to any detector built to look for it. The effect compounds further inside a chained pipeline, where one model drafts, another edits, and a third condenses or summarizes. Because KGW partitions are seeded by the preceding token, a word choice forced by an upstream model's green-list bias reshapes the partition that the next model in the chain will see for its own next token. The downstream model then applies its own bias on top of an already-distorted distribution, and these biases do not cancel each other out; they stack. A team that believes it runs a genuinely mixed workflow, blending multiple tools or models, may still be producing a detectable single-key fingerprint if one model dominates the token-level decisions early in the chain. The fingerprint can be read by parties the team never intended to inform, a compliance exposure whose full weight belongs with the detection standard itself.
Where the z-score threshold fails as a forensic standard
The z-score threshold that gives KGW detection its appearance of rigor is considerably less dependable than its widespread adoption suggests, and the scheme's own forensic reliability has been directly challenged. A 2026 empirical forensic evaluation documented what its authors call "paradoxical score increases": cases in which editing a watermarked text caused its z-score to rise rather than fall. The practical consequence runs directly counter to intuition. A content team that partially edits an AI-generated draft, under the reasonable assumption that human revision should make the text look less machine-produced, may instead end up with a document that scores as more detectably AI-generated than the unedited original. The same evaluation concluded that watermark detection results do not currently meet forensic readiness standards and cannot serve as reliable evidence of AI provenance in legal or compliance settings without additional corroboration.
Kirchenbauer and colleagues themselves recommend a high z-value threshold, above 4, as the bar for confident detection. That recommendation was built for a world in which text arrives unedited and whole. It was not built for a world of partial edits, chained models, and paraphrasing, and the paradoxical score increases documented in 2026 show what happens when that assumption breaks: a structural limit of the current partition-and-bias framework rather than a bug awaiting a fix, one baked into the shape of the framework itself, so any claim that a single z-score threshold settles the question of provenance has to be read against that structural limit. A 2025 paper presented at the ICLR GenAI Watermarking Workshop (arXiv:2505.23814) frames this as a three-way impossibility, with robustness against evasion, text quality preservation, and detection reliability in irreducible tension, such that improving any one degrades the others.
Sources
- A Linguistics-Aware LLM Watermarking via Syntactic Predictability
- Preprint. Optimizing watermarks for large language models Bram Wouters 1
- Robustness Assessment and Enhancement of Text Watermarking for Google’s SynthID
- Published as a conference paper at ICLR 2025 CAN WATERMARKS BE USED
- A Watermark for Large Language Models
- AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation