← All posts

How KGW Watermarking Works in Large Language Models

A lightweight technique embeds statistical traces in generated text that survive casual editing.

September 27, 2026

A watermark planted inside a language model's token probabilities does not sit on top of the finished text, it is baked into the math that produced the text, and that is why light editing cannot remove it. KGW watermarking works by splitting a model's vocabulary into two lists at every generation step, then nudging the model toward one list just enough to leave a statistical trace behind. Scrubbing that trace from finished prose is not a matter of trying harder with a thesaurus. A single model, watermarked or not, writes with a signature that survives paraphrase and gives way only to a fundamentally different kind of pipeline, one that never lets a single model's bias dominate the page. That is the argument this piece makes, and the rest of it is the evidence.

KGW watermarking and its origins

The scheme takes its name from its authors, Kirchenbauer, Geiping, Wen, Katz, Miers, and Goldstein, who published it at ICML in 2023. It has become the reference point against which nearly every later scheme gets measured, the null hypothesis of the field. The idea predates the paper, though: Scott Aaronson, working at OpenAI at the time, floated statistical watermarking in November 2022, months before KGW's formal scheme existed. Credit for the idea and credit for the mechanism belong to different people: Scott Aaronson, working at OpenAI at the time, floated statistical watermarking before KGW's formal scheme existed.

What the KGW authors built was deliberately lightweight. The watermark costs almost nothing in text quality, and detection needs no access to the model's API, its weights, or any other privileged information. Anyone holding the right key can check a piece of text without asking the model provider for cooperation, which is the design choice that made the scheme practical rather than merely clever.

KGW is a zero-bit scheme. It answers one question: was this text generated by a watermarked model, yes or no. It carries no user ID, no timestamp, no payload of any kind. Multi-bit schemes that embed richer metadata came later, and they solve a different problem.

How vocabulary partitioning works at generation time

The mechanism runs at the level of token generation, applied fresh at every step. A hash function takes the previous k tokens as input and uses that hash to split the model's vocabulary, call it V, into a green list of size γ|V| and a red list of size (1−γ)|V|. Under KGW's default settings, γ is 0.5, so roughly half the vocabulary counts as green at any given moment, though which half changes constantly depending on context.

The watermark itself is a bias term, δ, added directly to the logits of every green-list token before sampling. Raise the odds of an arbitrary, hash-selected half of the vocabulary, just a little, over and over, across an entire generation: that is the whole mechanism. No edit pass runs afterward, no substitution step gets applied. The bias lives inside the same math that decides what word comes next.

Turning δ up produces a stronger, more detectable watermark, but it distorts the text more, pushing the model away from what it would have said on its own. Turning δ down preserves quality but weakens the signal. Nobody gets both, and that tradeoff is baked into the math. It is baked into the math.

The k-gram context window at its limits

The choice of k, how many prior tokens feed the hash, controls a second tradeoff, separate from δ. A larger k generates more unique contexts, which makes the green list harder for an outsider to reverse-engineer. A smaller k means more overlap between contexts, so a local edit, swapping one word for another, is less likely to knock the watermark loose. Secrecy wants a large k. Robustness to editing wants a small one. No single value serves both goals, and that split is a limitation of the concept itself, not of any one team's tuning choices.

Pushing k all the way to zero produces UNIGRAM, the limiting case where the green list stops depending on context and stays fixed across the entire vocabulary for every position in the text. That fixed list is extremely durable: cut a sentence in half, paste in new text, and the watermark barely notices. But a list that never moves is also a list that sits still long enough to be mapped. Research has shown that fixed green lists can be mapped precisely because of that fixity, a vulnerability that static schemes like UNIGRAM cannot avoid. UNIGRAM does not escape the secrecy-robustness tradeoff, it just picks a point at the far end of it, trading one weakness for the other.

How detection works: the z-score

Detection reduces to a single statistical test. The formula is z = (|s|G − γT) / √(Tγ(1−γ)), where |s|G counts the green tokens actually found in a piece of text, T is the total token count, and γ is the green list's share of the vocabulary. What it measures is how far the observed green-token rate strays from what pure chance would produce. A high z-score means green tokens turned up far more often than random sampling would predict, a pattern traceable directly to the logit bias's effect on token frequency.

A z-score above 4 is widely used as a cutoff for flagging text as watermarked. That number matters well beyond KGW's own papers: it is the threshold attack researchers target when they test whether a scrubbing method actually works, and it resurfaces later in this piece as the line an entirely different kind of system has to cross.

KGW's original paper reported a false positive rate below 3×10⁻³% and a false negative rate below 1% under initial testing. Those numbers look close to airtight, and under the conditions they were measured in, they hold up. But lab conditions and the open internet are not the same environment, and the gap between the two is what the rest of this piece is about.

Why a single model's output is structurally traceable

The green-list bias is applied at every step of the generation, not once and left alone. It applies at every token position, across the full length of a generation, shaping the probability distribution assigned to every word from the first to the last.

A model working under one watermarking key produces a consistent statistical fingerprint across everything it writes, and that fingerprint does not fade as the text grows longer. It accumulates. The same signature repeats across every document that model produces under that key. Most people assume length dilutes a pattern. Here it concentrates one.

Research has found something uncomfortable for anyone hoping to keep watermarking quiet: hiding the fact that a watermark is in use at all is not straightforward, since watermark presence can potentially be inferred through black-box probing, not just KGW. The watermark's existence resists concealment as stubbornly as its content does.

That is why editing watermarked text after the fact does not solve the underlying problem. Swapping in synonyms, tightening a sentence or two, paraphrasing a paragraph, all of it will drop the z-score somewhat. But the logit bias that shaped the original token distribution does not get undone by editing around its edges. Lowering the score enough to matter requires replacing so much of the text that the edit becomes a rewrite, a different operation carrying a different cost.

KGW's place in the broader watermarking landscape, including production deployments

KGW functions as the root of a wider family tree. Later schemes branch off by optimizing for one of roughly five different objectives, and reading any comparison between schemes requires understanding KGW's tradeoffs first, since that is the baseline everything else gets measured against.

SynthID-Text is a production system, and it takes a genuinely different approach from KGW's logit bias, using Tournament Sampling instead. It never touches the model's training, only the sampling procedure applied at inference, and detection under this scheme is cheap and requires no access to the underlying model.

SynthID-Text runs inside the Gemini App and on the web. Feedback gathered across nearly 20 million watermarked and unwatermarked responses showed no statistically significant difference in how users rated response quality. That is a meaningful data point: watermarking at scale, in a shipped product, with no detectable quality tax attached.

A separate scheme, MORPHMARK, adapts the strength of the green-list bias to the model's cumulative probability mass sitting on the green list at that step. It addresses a known soft spot in logit-bias schemes: the difficulty with low-entropy text, where the model has so few reasonable next-token choices that biasing toward green forces obviously unnatural word choices.

Limits of KGW-based detection revealed by the attack research

Adversarial work against watermarking splits into two goals. Scrubbing tries to remove a real watermark, spoofing tries to forge one onto text that was never watermarked. Both matter, but scrubbing is where the empirical results are sharpest, and scrubbing, not spoofing, is what actually breaks KGW in practice.

Paraphrasing is the most direct scrubbing method, since a semantically equivalent rewrite disturbs the green-token distribution the original generation left behind. SynthID-Text itself has been shown vulnerable to meaning-preserving attacks of exactly this kind, including paraphrasing, copy-paste style modification, and back-translation. None of that makes the scheme worthless, but it does mean robustness claims need reading alongside the attacks actually tried against them, not just the ones the original paper anticipated.

BIRA, a query-free, black-box rewriting method, suppresses green tokens while keeping the meaning of the text intact. On a documented Llama-3.1-8B example, it took a z-score of 10.33 down to 2.60, clearing below the standard detection threshold of 4, with an evasion rate above 99%, clearing the bar with room to spare. That is not a marginal result. It clears the bar with room to spare.

The math behind why this works has a clean shape to it. If an attack manages to suppress green-token probability by a margin δ greater than zero below the induced threshold, the probability of successful detection decays exponentially in δ squared. Scrubbing is not a game of chipping away at the margins. Once an attack clears a certain threshold of suppression, detection does not degrade, it collapses. A light synonym pass tends to fail where a genuine rewrite succeeds, and there is not much middle ground between the two.

Multi-model authorship as the architectural response in Lettershred

Putting the pieces together, the diagnosis is structural. A single model applying KGW, or any single-key logit-bias scheme, produces a consistent overrepresentation of green tokens in everything it writes. That is the mechanism doing what it was built to do, and it is why the output stays traceable: the same bias that makes a watermark detectable is what makes a single model's output legible as a single model's output, watermarked or not.

Editing after the fact treats a structural condition like a surface stain. Paraphrasing, synonym swaps, light rewriting, all of it can push the z-score down by replacing individual tokens, but none of it disperses the logit bias that shaped the original distribution. It is a patch applied to a symptom while the cause sits somewhere the patch never reaches.

The alternative is a different architecture entirely: real dispersal of authorship across multiple models from different providers, mixed at the token level, so no single model's green-list signature ever gets the room to dominate the finished text. Under that setup, a z-score falls below the detection threshold because no single, consistent bias was ever applied across the whole document to begin with. Preventing a signature from forming beats trying to erase one after the fact, and that difference is the whole argument.

Lettershred's pipeline runs on exactly that principle. A brief goes in, a dozen model kernels draft their own versions in parallel, a judge panel picks the strongest line among them sentence by sentence, and a final enhance pass unifies the result into one coherent voice. The measured effect: a single-model baseline z-score of 15.46 drops to -0.32, well under the detection threshold of 4, with 177 of 240 tokens rewritten through that cross-provider dispersal. The watermark does not get erased after the fact. The watermark does not get erased after the fact; it never forms.

AI Watermark Detection

Sources