APIs, integration & security — in depth

Statistical Fingerprinting Methods Across Major Language Models

Every language model leaves a detectable statistical signature in its token choices.

Staff Writer · · 11 min read
Cover illustration for “Statistical Fingerprinting Methods Across Major Language Models”
LLM Fingerprinting · October 3, 2026 · 11 min read · 2,499 words

Every major language model leaves a statistical trace in the text it produces, and that trace is detectable because token sampling is never truly random. This piece traces how that fingerprint gets built, measured, and undone, along with the points at which it gives way.

How token sampling behavior produces a detectable statistical signature

A language model choosing its next word is not flipping a coin. At every step of generation, the model computes a set of scores, called logits, across its entire vocabulary, and samples the next token from the probability distribution those scores produce. That distribution is shaped by the model's training, its temperature setting, its repetition penalties, and a handful of other decoding parameters, all of which leave fingerprints in the pattern of choices a model makes over the course of a long passage.

The KGW scheme, introduced by Kirchenbauer et al., turns this fact into a deliberate signal rather than an incidental one. A fixed bias, denoted δ, gets added to the logits of every green-list token before sampling occurs, which tilts the odds of selection in their favor without necessarily forcing any single choice. Over the span of a full paragraph or essay, that small repeated nudge compounds into something measurable: text generated this way contains an abnormally high share of green-list tokens, a pattern that persists because the hash function re-runs at every single step of decoding.

What makes this mechanism worth understanding in depth is that it formalizes something that happens anyway. Any model with consistent sampling habits, whatever its temperature or nucleus threshold, produces a token-distribution regularity that a sufficiently patient detector can exploit, watermark or no watermark. KGW simply makes that regularity deliberate, controllable, and strong enough to detect reliably. That control comes at a cost built into the scheme from the outset: a larger δ produces a louder, easier-to-catch signal, while a smaller δ preserves the natural quality of the text at the expense of detection strength. Nearly every design decision in watermarking research traces back to this trade-off between visibility and fidelity.

Turning a Green-Token Bias into a Formal Accusation

Detecting a KGW watermark is a hypothesis test, structured the same way a drug trial or a clinical diagnostic is structured, with a null hypothesis, a test statistic, and a threshold for rejecting that null hypothesis with a stated degree of confidence.

The detector counts how many tokens in a suspect passage fall on the green list, then compares that count to what chance alone would predict. The formula is z = (|s|_G − γT) / √(γ(1−γ)T), where |s|_G is the observed green-token count, T is the number of eligible tokens in the passage, and γ is the expected fraction of green tokens if no watermark were present at all. Under the null hypothesis, meaning ordinary human writing or text from a model with no watermark applied, green-token counts follow a binomial distribution, the same distribution that governs coin flips and election polling samples. A z-score that lands far out on the tail of that distribution is statistically implausible without a watermark's intervention, and it is that implausibility, not any visual or stylistic cue, that triggers a positive finding.

The threshold for that finding is a dial that can be adjusted, not a fixed law of nature. Set low, the detector catches more true watermarks but also flags more innocent human writing as suspect. Set it high, and false positives drop, but so does the detector's ability to catch watermarked text that has been lightly edited or only weakly signaled to begin with. Production systems have to choose a point on that curve, and the choice carries consequences for anyone whose writing gets evaluated against it.

Real documents rarely arrive as uniform blocks of either human or machine text, and that reality complicates detection further. Mixed-source documents introduce a localization problem: the Geometric Cover Detector (GCD) and Adaptive Online Locator (AOL) extend watermark detection to mixed-source documents, GCD classifies whether any watermarked span exists, while AOL pinpoints the precise span boundaries, not just whether the whole document is. That distinction matters for any long-form content where a human editor and a model have worked on the same draft.

An assumption does most of the load-bearing work in the formula: that eligible tokens are drawn independently of one another, as the binomial model requires. The binomial calibration is the detector's Achilles heel as well as its strength: it assumes eligible tokens are independently drawn, an assumption that breaks down under low-entropy generation or when an adversary selectively replaces tokens. Both of those conditions deserve their own treatment, because they represent the two main fault lines running through every green-list watermarking scheme in production today.

Where the KGW fingerprint breaks down first

Watermarking depends on the model having a real choice to bias at each step. When the next token is nearly forced by context, that choice disappears, and the watermark has nothing left to act on.

Consider the prompt "import numpy as." Any competent language model completes that line with "np" essentially every time, since the entropy of that next-token distribution measures only about 0.048. If the hash function happens to place "np" on the red list at that position, the model faces an unpleasant choice: pick a green-list token instead and produce code that is grammatically or syntactically wrong, or keep "np" and let the watermark signal go uncollected at that position. Neither outcome serves the scheme well. The first corrupts output quality, and the second simply skips the opportunity to embed a signal.

The reverse case is just as troublesome. When the near-certain token happens to land on the green list by chance, the green-token count inflates without any real watermarking effort behind it, since the model was always going to pick that token regardless of bias. A detector reading that passage sees a strong z-score that originated from the text's inherent predictability rather than from any deliberate signal, and that confusion raises the risk of flagging ordinary human-written technical or formulaic prose as watermarked when it never was.

Researchers have built several responses to this specific weakness. Low-entropy POS-guided partitioning identifies parts of speech that tend toward high determinacy, function words, proper nouns locked into fixed phrases, code tokens, and exempts them from green-list biasing altogether, which concentrates the watermark's effort on positions where the model genuinely has room to choose. SWEET, from Lee et al., and EWD, from Lu et al., take a related approach: they watermark only high-entropy tokens, or give high-entropy tokens more weight during detection, though both require the detector to have access to the original model in order to calculate entropy in the first place, which introduces its own risk of model leakage. A method called Invisible Entropy removes that dependency by training a lightweight feature extractor and entropy tagger that predicts token entropy without needing the original model at all, and it matches state-of-the-art detection performance on the HumanEval and MBPP benchmarks while reducing the parameter overhead involved.

The upshot carries real consequences for anyone producing watermarked content in constrained genres. Source code, legal boilerplate, and heavily templated marketing copy are all low-entropy by construction, and the fingerprint embedded in text like this is structurally weaker, because the model has comparatively little room to express a preference at each step.

Context-Dependent Hashing and the Token Sequence

The green-red partition at any given position is recomputed fresh at every decoding step from a hash of the token or tokens that came immediately before, so the fingerprint belongs to the specific sequence of tokens produced, not to the model in some abstract sense.

First, it is what makes the scheme secret in any meaningful way: without the hash key, a detector has no way of reconstructing which tokens were green at each position in the sequence, so the key itself functions as the actual proof of provenance. Second, and more consequentially for anyone trying to defeat the scheme, the fingerprint is fragile with respect to the sequence it's embedded in. Changing one token early in a passage shifts every green-red assignment downstream of it, because each partition depends on the one before it. A single substitution, even one that leaves the meaning of a sentence completely intact, can degrade a z-score rapidly.

That same fragility explains what happens when content gets assembled from more than one model. If two models generate different portions of a passage, each with a different hash key or with no watermark applied at all, the resulting sequence carries no consistent green-list history running through it. No single z-test can flag a passage like that as watermarked, because the underlying chain the test depends on was never continuous to begin with. A fingerprint built on sequential dependence cannot survive a break in that sequence.

Three fingerprinting paradigms beyond green-list watermarking

Green-list watermarking is one branch of a larger family of techniques for establishing where text or a model's behavior came from, and the other branches operate on entirely different layers of the problem, serving different actors with different access to the systems involved.

Weight-based fingerprinting works directly on a model's internal parameters rather than on its output text. The Intrinsic Fingerprint approach, from Yoon et al. in July 2025, applies statistical summarization techniques, standard deviation, normalization, Pearson correlation, across a model's attention parameter matrices, producing a characteristic fingerprint sequence that survives even aggressive continued training or model upcycling. ProFLingo, from Jin et al. in 2024, pursues a related strategy, using adversarial prompts constructed to exploit a model's unique decision boundaries. Both require direct access to model weights, which makes them tools for model owners investigating theft or unauthorized derivative use, rather than tools a third-party content auditor could apply to text found in the wild.

Semantically conditioned watermarking represents a newer line of work. A 2026 ICLR paper describes a scheme where, instead of relying on fixed query keys that tend to break under fine-tuning, a statistical watermarking signal gets diffused throughout a model's responses to any prompt drawn from a broad, predetermined semantic domain. The model owner queries that domain later to detect the fingerprint reliably. The limitation is built into the design: it requires the model owner's own access at generation time, which makes it useful for an organization auditing its own model's outputs but unavailable to a third party trying to attribute content it did not generate itself.

What ties these approaches to green-list watermarking, despite their different mechanisms and different user bases, is a shared dependency. Weight-based fingerprinting, behavioral black-box methods, and semantically conditioned watermarks all rely on a single model generating output in a consistent, characteristic way. Each technique is built to collapse that consistency into a verifiable identity, and each one assumes that identity has a single, stable source to point back to. A content pipeline that genuinely draws from multiple providers and multiple models in combination does not present a consistent behavioral surface for any of these three paradigms to lock onto, which is less a flaw in any one method than a reflection of what each was built to assume in the first place.

Paraphrase Attacks, High-Entropy Targeting, and Knowledge Distillation

The same structural properties that make a fingerprint detectable, its dependence on a specific token sequence and a specific model's sampling bias, also make it removable, and the strongest attacks exploit that dependence with precision rather than brute force.

Paraphrasing is the most direct route. Untargeted and targeted training data paraphrasing, abbreviated UP and TP, applied before distillation eliminates inherited watermarks thoroughly, and targeted paraphrasing goes a step further by integrating the inverse of the extracted watermark rules directly into the paraphrase model, actively countering the green-list bias rather than simply scrambling tokens at random. A 2026 empirical evaluation found that every KGW and Unigram watermark initially detected in a test set was removed entirely after paraphrasing across a diverse range of prompts, with conditional removal complete across all valid paraphrase runs in that analysis; SynthID-Text performed only marginally better under the same conditions. The attacker doesn't even need to understand the specific watermark scheme in play. Running content through an expert LLM paraphraser, a single pass through a system like ChatGPT, drops detection rates for every tested watermarking method below 0.3, because the paraphraser simply regenerates the content from its own distribution and breaks the hash chain the original watermark depended on.

A more surgical attack targets only the tokens where the watermark signal actually lives. SIRA, from Cheng et al. in 2025, achieves near-total watermark removal by selectively rewriting only high-entropy tokens, the exact positions where green-list bias has the most room to operate, while leaving low-entropy tokens untouched entirely, which preserves the semantic integrity of the passage while dismantling the statistical signal running through it.

Knowledge distillation introduces a third pathway, both for inheriting a watermark and for stripping it back out. A technique called watermark neutralization applies the inverse of a watermark's rules during a student model's decoding phase after distillation, removing the watermark the student inherited from its teacher model while preserving the knowledge transferred in that same process, achieving removal without sacrificing the student's capability.

Removal and spoofing are distinct problems, and spoofing is the more troubling of the two for anyone relying on watermark detection as proof of origin. The DITTO framework, from Hyeseon Ahn et al. in 2025, uses knowledge distillation to make a non-watermarked model mimic the statistical signature of a watermarked one, which lets an adversary frame entirely innocent, human-written output as AI-generated. Separately, unified spoofing attacks demonstrated by Yi et al. in 2025 have forged signatures against KGW, Unigram, and SynthID-Text simultaneously. A positive detection result, in other words, does not guarantee the text came from the model the signature points to. That gap breaks the evidentiary chain that regulators and enterprise governance frameworks need intact if they intend to rely on watermark detection as proof of anything.

Google's SynthID-Text and the limits of single-provider watermarking at scale

In a 2026 empirical evaluation, every initially-detected KGW and Unigram watermark was removed after paraphrasing, with conditional removal complete at 100% of valid paraphrase runs, and SynthID-Text fared only marginally better. It has also been named alongside KGW and Unigram as a target of unified spoofing attacks capable of forging its signature on text the model never produced.

That outcome matters because SynthID-Text was not an under-engineered or experimental system. It represents a serious, well-funded attempt to solve the exact problem this piece has traced from its root, that the sampling bias driving any green-list watermark lives in a token sequence that paraphrasing, high-entropy rewriting, and knowledge distillation can all disrupt. The vulnerability sits in the architecture every token-level watermark shares: a signal built from sequential, context-dependent bias will always be exposed to attacks that target that same sequence and context. No amount of engineering resource at any single provider changes that fact, because the fragility is a property of the category, not a gap in any one company's execution.

Sources

  1. Watermarking with Low-Entropy POS-Guided Token ...
  2. Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarking
  3. Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?

More in LLM Fingerprinting