How KGW Watermarking Encodes Tokens in LLM Output
A technical deep dive into how the KGW watermarking scheme invisibly marks AI-generated text.

Text resists watermarking in ways that images do not, and the reason comes down to how each medium carries information. An image spreads its signal across millions of pixels, each one negligible on its own, so a watermark can hide in statistical variation no eye will ever register. Text has no such slack. Every word carries meaning, every character gets read, and a reader (or a simple script) can catch anything that doesn't belong. The only place a watermark can live durably in text is inside the generation process: before the words exist, when the model is still choosing among candidate tokens and the probability distribution over those choices can be nudged without leaving a visible seam.
That constraint is the reason the KGW scheme (Kirchenbauer et al., 2023) works the way it does: by intervening at the moment of generation rather than marking text after the fact, it sidesteps the fragility that doomed earlier approaches. The next step is to look at precisely what that intervention does, token by token.
Context Hashing, Vocabulary Partition, and Logit Bias at Each Token Position
KGW performs three operations at every single token the model generates, and the watermark visible in a finished passage is nothing more than the accumulated result of those three operations repeated hundreds or thousands of times.
The first operation is context hashing. At token position t, KGW takes the text that came immediately before it and hashes that context into a pseudorandom seed. The hash is deterministic: feed it the same context and the same secret key, and it produces the same seed every time. But without the key, the seed is unpredictable to any outside observer, which is what keeps the scheme secure.
The second operation is vocabulary partition. That seed is used to divide the model's entire vocabulary into two groups: a green list, containing a fraction γ of all possible tokens, and a red list, containing the rest. This partition is not a fixed roster of "watermarked words." A token can sit on the green list after one context and the red list after another, because the split is regenerated from scratch at every position based on what preceded it. One position later, the preceding context has changed, the seed changes with it, and the green list is rebuilt from scratch, possibly placing "woods" on green this time and "forest" on red.
The third operation is logit bias. Before the model samples its next token, KGW adds a constant bonus, δ, to the logit score of every token on the green list. This shifts the probability distribution in favor of green tokens without altering the underlying grammar or fluency of the output; the model is still choosing from the same plausible continuations, just with its preferences nudged.
KGW offers two ways to apply that nudge. In practice, soft embedding is the preferred choice, because it preserves the diversity of the model's output.
What makes this architecture resistant to casual detection is the same thing that makes it resistant to casual removal: the green list cannot be reconstructed by looking at a single token in isolation. Recovering it requires rerunning the entire hashing protocol, with the correct key, at every position in sequence.
Why the signal is invisible to readers but detectable by a statistical test
Invisibility and detectability come from the same source, which is what makes KGW usable. A logit bias small enough to leave any single sentence looking completely ordinary still accumulates, across hundreds of token positions, into a pattern no random passage would produce.
At any one position, the model usually has several plausible words it might choose next. A small bias toward the green list might tip the choice from one synonym to another, or favor one plausible continuation over an equally plausible alternative, without making that one sentence read strangely. One nudged word choice looks like nothing. Hundreds of nudged word choices, all biased in the same direction, become a measurable signal.
Detection works by reversing the generation process. A detector holding the same secret key retokenizes the candidate passage, reruns the context-hashing protocol at every position to reconstruct what the green list would have been at that point, and then counts how many of the actual tokens in the text landed on the green side.
A worked example makes the scale of the effect concrete. Suppose γ is set to 0.5, so half the vocabulary is green at any given position, and consider a passage 400 tokens long. Under the null hypothesis, that is, if the text were generated with no watermark at all, the expected count of green tokens would be 200, with a standard deviation of 10. If the detector instead counts 250 green tokens in that passage, the resulting z-score is 5, a result so far outside the expected range under random chance that it leaves little doubt the text was generated with the watermark active.
Detecting this signal requires no access to the model itself, only the secret key and the value of γ, both of which the deploying organization controls. That means detection can run as a lightweight service separate from the model, without exposing model weights or requiring heavy compute at inference time. Longer passages give the detector more token positions to work with, and confidence rises as length increases, though not in direct proportion: doubling the length of a passage does not double its z-score, even though it does push the score higher.
The four parameters that govern the watermark's behavior and their mutual tensions
Everything about how KGW behaves in practice comes down to four parameters: γ, the green-list fraction; δ, the logit bias; T, the usable length of the passage being tested; and h, the width of the context window used to seed the hash. An operator deploying KGW controls all four, but none of them can be tuned in isolation, because each one pulls on detectability, output quality, and security at the same time.
Lowering γ shrinks the share of the vocabulary treated as green at each position. That raises the excess of green tokens the model is pushed toward, strengthening the detectable signal, but it also shrinks the pool of tokens the model is allowed to prefer, which can force more awkward or constrained word choices. Raising δ has a related effect: a larger logit bias makes green tokens more likely to be chosen, which generally improves detection confidence, but it also distorts the model's natural distribution over words more aggressively, with a corresponding cost to how natural the output reads.
T, the number of token positions available to the detector, isn't something the operator tunes so much as something imposed by the length of whatever text is being checked. Short passages simply don't give the statistical test much to work with, while longer passages accumulate more evidence and support more confident detection.
The fourth parameter, h, governs how many preceding tokens feed into the context hash at each step. A wide context window creates a much larger space of possible states, making the partition harder to infer or reverse-engineer, but it also means a single substitution early in the text can cascade, shifting the green-list assignments at several downstream positions at once.
These four levers interact. A deployer who raises δ to guarantee reliable detection on short snippets of text pays for it in output quality. A deployer who widens h to make the scheme harder to attack makes that same scheme more brittle in the face of paraphrasing. There's no configuration of γ, δ, T, and h that maximizes detection confidence, output quality, and resistance to tampering all at once.
The detectability-quality-robustness trade-off that no current scheme escapes
Pushing watermark strength high enough to guarantee detection runs directly against the quality of the generated text, and this isn't a shortcoming unique to KGW's particular implementation. Research using the WaterMax framework, presented at NeurIPS 2024, identified this as a structural property of the entire family of watermarking approaches, describing it as a fundamental detectability-robustness-quality trade-off: no scheme now in use optimizes all three properties simultaneously.
The cost to quality doesn't spread evenly across a passage. This unevenness is also why perplexity, the standard measure of how fluent or natural a passage of text sounds, makes a poor benchmark for watermarked content: the fluency and coherence of a passage can remain largely unchanged even when the watermark has quietly corrupted something that matters semantically. Task-specific accuracy benchmarks, such as MMLU or GSM8K, capture that kind of semantic degradation far better than perplexity scores do, because they test whether the content is actually correct.
Part of the reason quality erodes unevenly traces back to how KGW applies its bias: uniformly, across every token position, regardless of how much freedom the model actually has at that position. A highly constrained slot, where only one or two tokens make grammatical or factual sense, pays a much higher semantic price for being pushed toward the green list than a loosely constrained slot with a dozen reasonable options. One correction to this problem, published at EMNLP 2025, introduced part-of-speech-guided token partitioning paired with a z-score-driven dynamic bias, and it achieved a measurable improvement in semantic fidelity over the standard KGW scheme. A related approach, the token-specific watermarking method presented at ICML 2024, adjusts both the green-list ratio γ and the logit bias δ per token, using lightweight learned networks conditioned on the representation of the preceding token, optimizing for detectability and semantic coherence together through multi-objective optimization rather than applying one fixed global setting everywhere.
Both of these refinements point to the same underlying diagnosis: the quality cost KGW imposes is a direct, structural consequence of applying one uniform γ and one uniform δ at token positions that are not uniform at all in how much freedom they actually offer the model. That cost gets baked into the text at the moment of generation, before any detector or any attacker ever touches it, which matters for what comes next: anyone trying to strip or exploit the watermark is working with a passage whose weak points were already determined when it was written.
The context-hashing chain creates a security-robustness tension the basic scheme cannot resolve
The same property that makes KGW's watermark hard to forge without the secret key is also what makes it fragile against certain kinds of edits: because each green list is derived from the context that came before it, any change to that context changes every green-list assignment downstream of the edit.
With a narrow context window, say h=1, where the hash seeds from only the single preceding token, a substitution anywhere in the text affects only the partition at the very next position. The trade-off is that a context derived from just one preceding token offers a much smaller space of possible states, which gives an attacker fewer unknowns to work against.
Widening the context window brings the opposite trade. A larger state space makes the partition far harder for an outside party to infer or reverse-engineer, which is good for security. But it also means a single word changed early in a passage can ripple forward, altering the green-list assignment at several subsequent positions, and that cascade can knock down the detectable green-token count well below what the original watermarked text would have produced.
This basic scheme carries a further weakness beyond edit sensitivity: it offers no guarantee against forgery. A detector that can be queried repeatedly by an outside party risks becoming an oracle, letting an attacker iteratively construct text engineered to trigger a positive detection result even though the model never generated it, which opens the door to spoofing. Later proposals, including alignment-based detection, semantic keys, and repeated-context masking, attempt to close these gaps, but each introduces a trade-off of its own.
Paraphrasing, API Exploitation, and Spoofing Attacks
Because the KGW watermark lives entirely in token-level probabilities rather than in any explicit marker, the ways to attack it follow directly from the mechanism itself, and the most effective methods work by exploiting that mechanism's own logic.
Paraphrasing is the most direct route. A more targeted method, the BIRA attack (Bias-Inversion Rewriting Attack, accepted at ICML 2026), goes further by applying a negative logit bias to a proxy suppression set identified through token surprisal, deliberately steering the rewritten text away from the tokens most likely to have been favored by the watermark. Tested on one language model, plain paraphrasing achieved only a moderate success rate at defeating detection on average, while BIRA reached a dramatically higher success rate, showing that attacks built around the watermark's own logic outperform generic rewriting by a wide margin.
Spoofing attacks work in the opposite direction, manufacturing a false positive. In what's called a piggyback spoofing attack, an attacker generates incorrect or harmful content engineered to match the z-score profile of genuinely watermarked text, making it appear as though a particular model produced content it never actually generated, with the potential to implicate the model's provider.
Cross-lingual translation and model distillation represent a third category of exposure. Translating a passage into another language replaces nearly the entire token sequence, destroying the context-derived partitions the original watermark depended on. Pan et al. Each of these routes exploits the same underlying fact: the watermark's signal exists only in the specific sequence of tokens chosen at generation time, and anything that substantially rewrites that sequence, whether a paraphrase, a translation, or a distilled imitation, threatens to take the signal down with it.


