← All posts

Watermark Persistence Across Translation and Summarization

Standard AI watermarks vanish when text is translated or summarized.

October 9, 2026

Cover illustration for “Watermark Persistence Across Translation and Summarization”

Whether watermarks persist across translation and summarization is a narrower question than it first sounds, and an entire regulatory framework's enforceability rests on the answer. The dominant text watermarking method used by large language models today produces a signal that survives editing, paraphrase, and even moderate rewriting, but collapses in specific, predictable ways once text crosses a language boundary or gets compressed. Understanding why that collapse happens, and why it happens differently for translation than for summarization, is the clearest way to judge how much any single watermark can actually be trusted.

How the KGW green-list scheme encodes a detectable signal in every token

The scheme behind most deployed text watermarks today traces back to Kirchenbauer and colleagues, commonly shortened to KGW. At each step where the model picks the next word, a hash of the preceding tokens splits the entire vocabulary into two groups: a green list and a red list. Before the model samples its next token, a fixed bias, called delta, gets added to the score of every green-list candidate. The model still produces fluent, ordinary-looking text, but it leans toward the green list slightly more often than chance would predict.

Detecting the watermark later means re-running that same hash-based split on a piece of text and counting how many of its tokens landed in the green list. That count feeds a z-test: the observed green-token fraction gets compared against the fraction you'd expect under pure chance, and the resulting z-score grows as the text gets longer. A short passage gives the detector little to work with. A long one gives the statistical test real power, because the z-score scales with the square root of the token count.

What matters most for everything that follows is that this signal is lexical. It lives in which specific words the model chose, measured against a specific vocabulary's partition into green and red. Google DeepMind's SynthID-Text, which runs on Gemini, uses a different detection formula (a score built from mean or Bayesian estimates over what are called g-values, rather than a straightforward green-ratio z-test), but it still belongs to the same red-green family, and outside researchers have shown a red-green z-test can pick it up as a black-box probe. A separate scheme called Stateless Bernoulli Watermarking, built by Ceppi and Sanchez at the European Commission's Joint Research Centre, swaps out KGW's vocabulary permutation for independent coin-flip trials on each token, and still preserves the same z-score detection guarantee. However the membership test gets computed, the z-test framework underneath is the durable part of the design. And all of these schemes share the same constraint: push delta higher to make the signal easier to detect, and the text quality degrades enough that very high delta values are not practical to deploy. The tension between signal strength and output quality exists before anyone tries to remove the watermark.

Translation Destroys the Watermark Signal

The green list a detector checks against was built from one language's vocabulary. Translate the text into another language, and the tokens being checked no longer belong to that vocabulary at all, so the detector has nothing coherent left to measure. Researchers call this the cross-lingual watermark removal attack. It is the cheapest documented way to get AI-generated text past a detector, and it does not require any special skill or adversarial intent.

The data backs this up directly: the watermark signal measured in a piece of text barely correlates with the signal measured in that same text's translation. The green-token fraction on the translated side reverts toward what you'd expect from chance, because the tokens being sampled now come from a different vocabulary with a different key mapping, or often no mapping. A multinational content team running AI-drafted copy through a standard machine translation layer strips the compliance marker as an incidental side effect of normal work, not as a deliberate evasion. Follow-up research from 2025 found that even watermarking schemes purpose-built for cross-lingual robustness hold up only in a handful of high-resource languages, and the exposure grows as content moves toward smaller-language markets.

One obvious counter is to anchor the watermark to meaning instead of to specific tokens, so that translation couldn't touch it. The XSIR scheme takes exactly this approach. Whether that semantic anchoring survives contact with real-world pipelines is a separate question, and the answer turns out to be no once summarization enters the picture, which is the subject of the next section. Translation on its own is close to a full removal of the signal; what summarization does to what's left is a different mechanism entirely, and the two compound.

Summarization Degrades the Watermark Through Length Compression

Summarization does not attack the watermark by replacing vocabulary. It attacks the statistical foundation the detector depends on: text length. Because the z-score scales with the square root of the token count, cutting a document's length in half reduces the maximum achievable signal by roughly 30 percent, even if every surviving token is still green. Aggressive summarization can compress a document down to a fraction of its original length, and the z-score falls below any fixed detection threshold well before the green-token fraction itself drops to zero.

A second effect compounds the length-driven loss: summarization tends to keep the words carrying the most meaning and cut the rest, and those high-salience words are exactly the tokens where the model had the least freedom to lean toward the green list in the first place, since the content demanded a specific word regardless of which list it fell on. The tokens that survive summarization are disproportionately the ones where the watermark's bias had the weakest grip to begin with.

The XSIR scheme, designed specifically for cross-lingual robustness, shows how much this costs in measured terms. Paraphrasing alone still leaves an AUROC of 0.827, and a cross-lingual pivot attack through another language leaves 0.823, both still clearly detectable and far from chance-level performance. Cross-lingual summarization, by contrast, drives AUROC down to 0.53, which is close to indistinguishable from a coin flip. That gap shows how these mechanisms differ: summarization alone is a partial attack, translation alone is close to a complete one, and running both together produces the most effective removal pathway documented so far, because the two mechanisms don't just add, they compound.

Tuning delta upward doesn't rescue the scheme here either. Research on optimizing watermark strength for large language models shows that a higher delta does make the signal somewhat more resilient to length compression, but it also degrades text quality enough to make the tradeoff impractical at the levels that would be needed. There's no way to dial the parameter up far enough to beat summarization without producing text nobody would want to publish. Mixed-source detection, where a watermarked segment sits inside a longer human-written document, is already difficult enough that researchers built specialized interval detectors like GCD and AOL just to handle it. Summarization makes that problem worse still, by blending and compressing the watermarked span until even those specialized tools have less to work with.

Cross-lingual summarization as the practical removal pathway

Generate content in one language, summarize it through a pivot, and publish it in another, and the result is cross-lingual summarization, or CLSA. That is the ordinary architecture of multinational AI content production, so watermark removal happens as a side effect of routine operations rather than as something an adversary has to engineer.

A generate-translate-summarize chain combines the two failure modes directly: translation replaces the vocabulary space the detector was built to check, and summarization dilutes the signal through length compression. Run both steps on XSIR and detection falls to near-chance, while the resulting summary remains coherent and genuinely useful to the reader it was written for. Nothing about the output looks broken or suspicious. This is how AI-assisted localization already works at scale: draft in a high-resource model, compress the draft for a target audience, translate it for publication in the market that needed it.

Content engineering teams running these pipelines have already noticed the pattern appears differently depending on the route content takes. An image path through the same pipeline may preserve its watermark intact, while the text-summarization path strips it after paraphrase. The vulnerability is route-dependent and invisible unless someone instruments each route separately and checks. The compliance consequence follows directly: a team that believes its AI-generated content is detectably watermarked can be wrong the moment that content passes through a completely standard localization step, with no way to know unless it tests the final, published output rather than trusting the generation step. The content team only sees the final text, and that text reads naturally with no visible trace of what got stripped. Only running a detector on the published output reveals the loss, and most teams are not doing that.

Multi-Model Pipelines and the Attribution Problem

A separate failure occurs once watermarked text passes through a second watermarked model, which happens in any chained workflow where both the generation step and the translation or summarization step run on LLMs. The second model's watermark overwrites the first, so provenance attribution breaks down as a structural consequence of the pipeline itself.

A NAACL 2025 study, "Lost in Overlap" by Luo and co-authors, measured this directly: rewriting watermarked text with a second watermarked model dropped detectability of the first mark to somewhere between 0.2 percent and 15 percent, while the second mark read out at 91 to 99 percent. A generate-then-translate workflow using an LLM for the translation step would attribute the resulting content to the translation model, not to the model that actually generated the original ideas, which is the opposite of what any compliance framework is trying to establish. A paper titled Watermarks Attack Watermarks (arXiv:2605.16796) formalizes this as a general pattern: re-watermarking functions as a removal attack on whatever watermark came before it, whether or not anyone intended it that way.

A related risk runs in the opposite direction and lasts far longer. Meta researchers, Tom Sander and colleagues, described this at NeurIPS 2024 as radioactivity: fine-tuning a model on a corpus where even a small fraction of the training text carries a watermark leaves a detectable trace in that new model's own behavior, with an extremely low rate of false alarms. For any content team feeding AI-generated text back into a training set, that creates a second provenance risk running alongside the first: a watermark that got fully stripped from a published, customer-facing document can still be sitting inside the training data that shapes the next model's behavior. Single-model pipelines carry a particular exposure here, since one model produces a repeatable statistical signature across every sentence it writes, and chaining that model to a second LLM for translation or summarization means the signature that eventually reaches a detector is not the one the original team meant to stand behind.

What the EU regulatory framework now requires in transformation pipelines

Regulators have already built a response to the fragility described above, and it creates its own gaps at exactly the pipeline steps where that fragility is worst. The EU Code of Practice on Transparency of AI-Generated Content was published in June 2026, with Commission Guidelines finalized July 20, 2026. Providers of AI systems that generate synthetic text content must mark their outputs in a machine-readable format, with the obligation applying to new systems starting August 2, 2026, and a grace period running through December 2, 2026 for systems that already existed, covering the Article 50(2) machine-readable marking obligation specifically. Non-compliance carries fines of up to €15 million or 3 percent of annual global turnover.

The draft Code does not treat any single watermarking technique as sufficient on its own to satisfy Article 50(2). It calls for a multilayered approach, combining at least two machine-readable methods, such as digitally signed metadata alongside imperceptible watermarking, with a single technique allowed only under two narrow exceptions. That requirement exists precisely because single-technique watermarking has the structural weaknesses described above. But the Act's own language around interoperability and robustness includes the qualifier "as far as technically feasible," and that phrase opens exactly the ambiguity this piece has been tracing: cross-lingual summarization and multi-model chaining are the pipeline steps where robustness is least achievable under any watermarking scheme that currently exists.

Around 190 signatories to the Code of Practice have committed to providing detection systems. OpenAI had already built a text watermark and held it back from release, before announcing deployment for EU users on October 5, 2026. Providers still working toward deployment carry the same legal obligations with no independent way for outsiders to verify their assurances in the meantime. Organizations that failed to capture provenance information at the moment of generation cannot reconstruct it after the fact. The EU AI Act requires documented provenance for high-risk AI systems, and the transformation failures described in the sections above mean provenance recorded at generation time may simply not survive to the version of the content that eventually reaches a regulator's desk. The regulatory body setting these rules is also building the infrastructure meant to satisfy them, an indication that the expectation of working technical solutions is built into the policy from the inside.

Architectural diversity across models as a structurally sounder response than patching a single-model watermark

Every attempt to patch the KGW family, whether by raising delta, anchoring the signal to meaning instead of tokens, or building cross-lingual robustness the way XSIR does, still leaves behind a single model's repeatable statistical fingerprint. Translation and summarization can remove that fingerprint because the fix never touches the actual point of failure: one model, one token distribution, one detectable signature standing in for the whole pipeline. A pipeline built on the assumption that a single watermark will carry proof of provenance through every translation, summary, and model handoff is relying on a guarantee the underlying statistics were never built to provide. The more durable response is architectural, treating provenance as something that has to be tracked across every model and every transformation step in a pipeline, rather than something one watermark at the start can be expected to carry to the end on its own.

AI Watermark Detection

Sources