← All posts

Open-Source vs. Proprietary AI Watermark Detectors Compared

Open-source detectors offer auditability but fall short as legal evidence.

October 6, 2026

Cover illustration for “Open-Source vs. Proprietary AI Watermark Detectors Compared”

Nearly every detector a content team will encounter traces back to one architecture: the KGW scheme, first described by Kirchenbauer and colleagues in 2023. Any serious comparison of watermark detectors has to start here, because the vocabulary it introduces (green list, logit bias, z-score, detection threshold) governs how almost every tool on the market, open or proprietary, makes its case. Over a long enough stretch of text, that nudge adds up to something statistically detectable: more green tokens than a random, unwatermarked document would ever contain.

Detection works by running a one-sided z-test on that green-token excess. The z-score counts how many standard deviations the observed green-token rate sits above what you'd expect from an unwatermarked baseline. A score that falls below the conventional detection threshold of 4 signals that it isn't, or that whatever watermark was there no longer survives in a form the detector can find. That single number, the z-score against a threshold, is what every claim about watermark presence or absence ultimately rests on, whether a tool reports it to the user or hides it behind a label.

Two extensions build on that same core, including the Adaptive Online Locator (AOL), which goes a step further, trying to find where that segment starts and stops. None of this says anything yet about which tools implement it well, who controls the key, or how easy the whole apparatus is to defeat. Those questions come next, but they only make sense once the z-score itself is understood as the operational unit of measurement that the rest of this argument turns on.

What the open-source and proprietary camps look like in mid-2026

With the mechanics in place, the next question is who has actually built on them, and on what terms. Open-source watermarking schemes available for independent deployment include KGW itself, Unigram, GaussMark (Block et al., 2025), and Unremovable (Christ et al., 2024). Any of these can be tested by a content team without needing permission or cooperation from the model vendor that produced the text.

The proprietary side looks different, and the differences run deeper than cost. Google DeepMind put SynthID-Text into production on Gemini starting in May 2024, published the underlying method in Nature, and released the scheme itself on Hugging Face. Anthropic began watermarking all new Claude models starting August 2, 2026, using a version of Google DeepMind's SynthID-Text statistical sampling approach. OpenAI has no deployed text watermark at all, only concept-stage work that hasn't reached production.

The structural asymmetry that falls out of this is the real finding of this section. Proprietary schemes lean on key secrecy as their main defense: if nobody outside the vendor can see how detection works, nobody outside the vendor can reverse-engineer a bypass either. But that secrecy cuts both ways. A third party, whether a journalist, an auditor, or opposing counsel in a dispute, cannot independently confirm a proprietary detector's finding without the vendor's cooperation, and right now that cooperation is either waitlisted (Google) or not yet available (Anthropic) or simply doesn't exist (OpenAI).

Regulatory Deadlines Make the Open-Source vs. Proprietary Choice Operational

Two deadlines have already passed, and each one turns the open-source versus proprietary question from a research debate into a documentation burden a compliance team has to carry on a specific date. The EU AI Act's Article 50(2) marking obligation requires that generative AI outputs be marked in a machine-readable format and detectable as artificially generated. California's AI Transparency Act set its own deadline earlier: providers above a large-scale monthly-user threshold had to have invisible watermarks in place by January 2026, under a disclosure standard that the statute itself describes as needing to be "permanent or extraordinarily difficult to remove."

Both rules point toward the same practical demand. A provider has to be able to show, at audit, that its labeling pipeline produces marks that are persistent, that hold up under ordinary use, and that an independent party can verify. Open-source schemes don't have that problem on their face: anyone can run the detection code, reproduce the test, and confirm or challenge the result, which lines up much better with the kind of independent reproducibility that evidentiary standards like Daubert expect from any technical method presented as proof.

That advantage for open-source schemes comes with a catch, though, and it's a serious one. Reproducibility only matters if the thing being reproduced is actually reliable once somebody tries to defeat it. That's the question the next section takes up directly.

The forensic readiness problem that undermines every current detector, open-source and proprietary alike

A study submitted to the AAAI/ACM Conference on AI, Ethics, and Society (AIES) 2026 tested three representative methods, KGW, Unigram, and MarkLLM's implementation of SynthID-Text, against the Daubert admissibility criteria that U.S. courts use to evaluate expert evidence, and against the NIST SP 800-86 digital forensic process that enterprise risk frameworks lean on. These are the actual standards a watermark would need to satisfy if its output were ever going to function as proof in a legal or compliance setting, and the results are sobering for anyone counting on watermarking to do that job today.

Even before anyone tried to attack the watermarks, baseline false-negative rates were already high across all three methods, higher for Unigram than for KGW or SynthID. Once a simple meaning-preserving paraphrase was applied, the picture got worse. Every piece of text that had initially been detected as watermarked under KGW and Unigram lost that watermark entirely after paraphrasing, a complete, 100% conditional removal rate. SynthID held up only slightly better, with removal reaching 98.3%. SynthID also showed problems in the other direction: it flagged a meaningful share of paraphrased human-written text as AI-generated, produced a substantial rate of paradoxical results, and placed the majority of its own untouched, genuinely watermarked output into an uncertainty deadband where the tool itself couldn't commit to a verdict.

None of the three methods satisfied more than two of the five Daubert factors that govern courtroom admissibility. The right way to read this finding is as a recalibration: watermarking functions today as a deterrent and a signal of intent, not as a forensic instrument a legal team can rely on to prove provenance after the fact. That reframing carries directly into the next question, which is how text actually loses its watermark, and why some architectures are more exposed to specific attacks than others.

The three evasion classes and which detector architecture is more exposed to each

Three distinct evasion classes explain how watermarks fail in practice, and each one exploits a different weak point depending on whether the detector is open-source or proprietary. A content team that can't tell these apart has no real way to judge what a detector's claims are actually worth.

Paraphrase attacks are the most legally realistic of the three, because a meaning-preserving rewrite is hard to call evidence tampering when the meaning is unchanged. Color-aware substitution attacks work differently. That kind of attack lands harder against open-source schemes, precisely because the detection logic sitting out in the open that makes independent audit possible also makes the green-list structure visible to anyone trying to reverse it.

Spoofing attacks run in the opposite direction from the first two. The DITTO framework (An et al., 2025) uses knowledge distillation to train an attacker's model to produce text that falsely appears to carry a target model's watermark, corrupting attribution. Existing watermarking schemes are built mainly to resist paraphrasing and largely leave spoofing unaddressed. A 2026 ACL Findings paper called DualGuard is described as the first algorithm to claim defense against both paraphrase and spoofing at once, striking a workable trade-off between the two without giving up detectability. Separately, researchers at Tsinghua University led by Pan and colleagues found that targeted training-data paraphrasing and inference-time watermark neutralization both thoroughly erase inherited watermarks in knowledge-distillation settings, the exact scenario where watermark "radioactivity" is supposed to flag unauthorized distillation, while untargeted training-data paraphrasing only removes the signal partially.

Open-source detectors are more vulnerable to color-aware substitution because their logic is public by design, while proprietary detectors with undisclosed keys resist that specific attack better but remain just as exposed to meaning-preserving paraphrase, which needs no knowledge of the key at all to succeed.

What transparency of scoring means for a content team evaluating detection risk

A detector that hands back a raw z-score gives a content team something to actually work with: a number they can benchmark, set thresholds against, and check again later. A detector that hands back only a label or a probability from a model nobody outside the vendor can inspect gives the team a verdict and nothing to push back against. Proprietary detectors that keep the scoring mechanics hidden leave a team unable to tell whether a positive flag came from a borderline score or an extreme one, what the false-positive rate looks like on their specific kind of content, or whether the underlying model has quietly changed since the last time they checked.

Understanding the z-score mechanics matters for a second reason beyond reading a detector's output: it's also what lets a content team evaluate workflows built specifically to reduce detection risk at the architecture level. The Lettershred benchmark offers a concrete illustration of what a reproducible claim in this space looks like: z-scores dropping from 15.46 to -0.32 against a KGW-family detector, tested with a named method, a stated threshold, and a stated result, which is the kind of evidence a team can bring to an auditor or a legal review because every part of the claim can be checked.

The LDiet detection method (Zhao et al., ICLR 2025) shows the same problem from the opposite angle. The practical question for a content team is what z-score that detector is using, what threshold it's measured against, and whether the pipeline that produced the content is built in a way that keeps it under that line.

How multi-model pipeline architecture changes the watermark exposure profile

A single-model content pipeline leaves the same repeatable statistical signature across every token it produces, and that signature is what KGW-family detectors exist to find. The exposure here is architectural, a function of how the pipeline is built rather than how carefully someone edits the output afterward. Each model carries its own green-list assignment, shaped by its own weights, decoding settings, and provider. When output from several models gets blended at the token level, rather than stitched together in separate chunks, the resulting green-token distribution spreads across providers and no longer matches the statistical fingerprint of any single model.

Cascaded pipelines, where one model's output feeds into another stage, have a side effect: each transition between stages, whether from one model to the next or from a model into a post-processing step, can disrupt the prior stage's token-level signal. That disruption is incidental, though, not a designed compliance outcome, and it doesn't produce a verifiable z-score result a team could point to with confidence. The strongest position in that literature is to layer all three mechanisms together, but even that layered approach doesn't touch the underlying problem of a single model's statistical fingerprint sitting inside the text.

Genuine multi-model architecture is built to address that underlying problem. The outcome is measurable: the vast majority of tokens get rewritten through that cross-provider process, and the z-score against a KGW-family detector drops from 15.46 to -0.32, well under the conventional detection threshold of 4. The structural asymmetry between open and proprietary detectors described earlier also matters here. Open detectors can be studied closely enough to guide evasion strategies, while proprietary detectors block that kind of independent study but remain just as vulnerable to a dispersal-based architecture like this one, since neither detector family depends on knowing how the text was produced to begin with.

Evaluating any detector, open-source or proprietary, against your pipeline's actual risk

The question for a content team is whether a given detector's scoring is transparent enough to benchmark against, and whether the pipeline producing the content is built in a way that stays under whatever threshold that detector applies. Everything else in this comparison folds into that one test. A detector worth trusting reports a raw z-score rather than a bare label, because a number can be checked, challenged, and reproduced, while a label from a hidden model can only be accepted or disputed on faith. It should also publish, or at least make discoverable, the detection threshold it applies and the conditions under which that threshold was set, since a threshold tuned for one content domain, as the LDiet work on IP-infringement detection shows, won't necessarily hold up in another.

A defensible evaluation also has to account for which evasion class the detector is actually exposed to. Given the forensic readiness findings, no detector in current use, open or proprietary, should be treated as courtroom-grade or audit-grade proof on its own; each should be understood as a deterrent signal that still needs to be backed by documentation a team controls independently. For a team that needs to show an auditor something concrete, whether under the EU AI Act's Article 50(2) marking obligation or California's permanence-and-difficulty-of-removal standard, the strongest position is a pipeline architecture that produces its own reproducible numbers rather than one that depends entirely on a vendor's detector staying available, accurate, and willing to cooperate. A multi-model blending pipeline like Lettershred generates exactly that kind of evidence, a named test, a stated threshold, and a measured result, that holds up regardless of which detector architecture a given compliance framework happens to recognize.

AI Watermark Detection

Sources