Cross-Provider Token Dispersal and Its Effect on Fingerprint Strength
Mixing text from different AI providers breaks watermark detection without editing a word.

Cross-provider token dispersal defeats KGW-family watermark detection without editing a single character of the finished text. The mechanism is structural: dispersal removes the statistical precondition the detector needs, rather than hiding or scrubbing the signal after the fact.
How the KGW z-test works
The KGW scheme, the baseline against which nearly all watermark and evasion research gets measured, works by biasing a language model's own token choices at the moment of generation. At each decoding step, a secret key combined with the recent context splits the vocabulary into two groups, a green set and a red set, and the model's logits for green tokens get a fixed boost, called δ. That boost makes the model lean toward green tokens more often than chance alone would produce, and that lean is the entire watermark. There is no hidden signature embedded in the text, no metadata tag riding along with it. It all comes down to a simple tilt in which token gets picked next.
Detection reverses the process with a one-sided z-test on exactly that tilt: Z_K(T):= (N_g − γn) / √(γ(1−γ)n), where N_g counts the green tokens actually present, γ is the green fraction you'd expect under pure chance, and n is the token count. A detector calls a document watermarked when Z clears a set threshold, z_0. The test is cheap to run and fits neatly into standard decoding pipelines, which is a large part of why it became the field's reference point. The bias δ is the dial that controls how strong the signal is: push it higher and the green-token excess becomes easier to detect, but the output drifts further from what the model would have said unprompted, a trade-off the design bakes in.
What makes this detail matter for everything that follows: the z-test only works if the detector can see an unbroken sequence of tokens long enough to recover the green-token bias, generated under one consistent key, at a reasonable decoding temperature, with enough length to let the statistics settle. If any one of those conditions is missing, the z-score drifts back toward the null, indistinguishable from unwatermarked prose. That dependency, on an intact, single-key, sufficiently long token stream, is the exact thing cross-provider dispersal interrupts.
Why the z-test fails across provider boundaries
When different segments of a document come from different providers, each operating under its own secret key, or under no watermark scheme at all, the green-token excess the z-test depends on never has a chance to build. A token can count as green under Provider A's key K_A, but it has no such status under Provider B's key K_B. A detector that applies a single key to the merged document ends up reading a blend: green-and-red tokens from A sitting alongside red-and-green tokens from B, measured against a key that only matches one of them, if it matches either. The green fraction that the test expects to see averaged across this mixed population collapses toward whatever a coin flip would produce.
The documented mechanism is the divergence between the actual green ratio in the merged text and the preset green ratio the detector assumes. That divergence breaks z-score detection under a fixed threshold outright, and it means a watermark may surface under one hash key while staying invisible under another, depending entirely on which provider's segments happen to dominate a given stretch of text. KGW's whole detection apparatus leans on token-distribution regularities that hold within a single generation run. Mixed-source documents break that regularity by design, and when a multi-provider content pipeline routes paragraphs or sections to different models as a matter of course, it reproduces this exact failure at a scale no lab experiment captures.
None of this involves touching a single token after generation. No one rewrites a sentence or swaps a word to throw off the detector. Dispersal is not an attack performed on finished text; it is an architectural condition that prevents a coherent signal from accumulating during generation.
A fair objection follows naturally: couldn't a long enough document, dominated by one provider, still accumulate enough green-token runs from that provider alone to clear the detection threshold, mixing or no mixing? The z-test needs an unbroken sequence long enough to reject the null hypothesis with confidence, so the objection holds if you assemble a document from one giant block per provider. But that is not how multi-model content pipelines tend to operate. A pipeline built around a judge-panel architecture, selecting the best line or the best paragraph from among several candidate models at each step, breaks the sequence at the sentence or paragraph level, long before any single provider's contribution reaches the length a z-test needs to accumulate signal. The interruption happens too early and too often for the statistic to recover.
Low-entropy decoding independently weakens the watermark before any mixing occurs
A second mechanism degrades the watermark before cross-provider mixing ever enters the picture, and it has nothing to do with how many providers are involved. Content pipelines routinely run factual or constrained sections, code, citations, structured data, at low decoding temperature to keep outputs accurate and repeatable. Low temperature sharpens the model's logit distribution so sharply that token choices concentrate on one or two dominant candidates, so the green-bias δ has almost nothing left to work with. The model is already near-deterministic about what comes next, and a fixed logit boost can't meaningfully redirect a choice the model had already all but made.
The paper "Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarking" (arXiv 2505.14112) formalizes exactly this failure. In low-entropy scenarios, where the next token is highly predictable regardless of which list it falls on, KGW runs into a bind on both sides. If the expected, highly predictable token happens to sit in the red list, the model either picks it anyway, which weakens the watermark's green-token excess, or gets forced onto a green-list substitute, which disrupts the fluency of the output. If the expected token happens to sit in the green list, the green-token count inflates but carries no real signal, because the model would have chosen that token anyway, regardless of the bias.
Low-entropy positions are common in professional content pipelines: code snippets, proper nouns, formulaic transitions, and list items are all near-deterministic, contributing noise to the eventual z-score. The Invisible Entropy paper's proposed fix, a lightweight feature extractor paired with an entropy tagger that predicts low-entropy positions and skips watermarking there, assumes a single model owner controlling the entire generation process end to end. That assumption fails for a team that routes work across multiple providers and multiple models, where no single party can see the full decoding process to apply such a fix consistently.
The weakened watermark signal and the degradation from cross-provider mixing compound each other. A pipeline that assigns technical or factual sections to one provider and runs them at low temperature has already weakened that provider's watermark contribution before a single token from another provider enters the document. Cross-provider mixing then degrades whatever signal survived that first pass. Dispersal, in other words, picked up an ally it never had to recruit.
Heavy watermark bias harms output quality, creating a quality–detectability bind for single-model pipelines
For a team committed to a single model and a single watermark key, there is one lever left to compensate for a weakening z-score: raise δ, the bias that pushes the model toward green tokens. Raising it does restore detectability, at least on paper. It also measurably distorts the text, because a larger forced preference for an arbitrary subset of the vocabulary pulls word choice away from what the model would naturally produce. It's a structural tension built into any watermarking scheme that relies on an accumulated statistical bias to produce a detectable signal: the stronger the bias, the stronger the signal, and the stronger the bias, the more the output reads like it was written under duress.
Research has tried to soften that trade-off. WaterMod (Park et al., AAAI 2026) uses modular token-rank partitioning to cut the distortion a given bias level produces, and it improves the quality side of the equation measurably. But the sources are clear that WaterMod leaves cross-provider dilution unresolved. It makes the single-model bind more livable without resolving the structural fragility the z-test has when it meets a mixed-source document.
For a content team, this bind is operational. Publishing visibly distorted text just to keep a watermark detectable is not something an editorial operation can sustain, since readers notice stilted phrasing and unnatural word choices long before a detector runs. Dial the bias back down to protect quality, and the z-score drifts toward the null, the same outcome dispersal produces, except arrived at voluntarily and at the direct cost of detectability the team was trying to preserve. Dispersal sidesteps the whole trade entirely: the z-score collapses as a structural consequence of the architecture, with no bias reduction required and no quality sacrificed to get there.
Editing and paraphrasing attacks are a different problem from dispersal
Watermark evasion through paraphrasing and rewriting is a long-studied, well-documented family of attacks that operates at a different layer than dispersal. A paraphrasing attack takes a finished, watermarked sequence and changes its tokens after generation, aiming to drop the green-token count below the detection threshold. Dispersal prevents the watermark from accumulating during generation. One is a cleanup operation performed on evidence that already exists; the other prevents the evidence from forming.
A paper from September 2025 sorts evasion attacks into query-based and query-free families. Query-free attacks work in a black-box setting directly on generated text: they typically use LLM-based rewriting that masks and regenerates high-entropy tokens, or fine-tune a dedicated paraphrase model to do the job. A separate paper out of Tsinghua on watermark robustness against knowledge distillation (arXiv 2502.11598) shows that targeted training-data paraphrasing and inference-time watermark neutralization can eliminate watermarks a distilled model inherited from its teacher. Both techniques carry real costs: watermark neutralization requires knowing the watermark's rules in advance, and training-data paraphrasing measurably degrades how well the student model learns from its teacher.
These are all post-hoc operations. They accept that a single model produced a detectable, watermarked sequence, and then spend compute, latency, and quality risk trying to erase what's already there. Dispersal leaves no such sequence to erase. Each provider contributes tokens under its own key, or under no key at all, and no operation performed on the merged document after the fact can reconstruct a single-model sequence that was never generated to begin with, because there is nothing to reconstruct.
The inverse of erasure, forging a watermark rather than removing one, shows how manipulable this layer of infrastructure really is. The DITTO attack, published in mid-October 2025 (arXiv 2510.10987), used watermark knowledge-distillation to train a surrogate model whose output became statistically indistinguishable from genuinely watermarked text, forging provenance. Evasion and forgery pull in opposite directions, but they share the same lesson: token-level watermarks function as fragile infrastructure whenever the generation architecture isn't controlled end to end, whether that fragility is being exploited deliberately or not.
Provider infrastructure variation means even single-model routing can inadvertently disperse fingerprints
Fingerprint instability can appear even without a deliberate multi-provider architecture. It can appear inside what looks, on paper, like a single-model pipeline, simply because of how model providers run their infrastructure. A 2026 study on LLM API gateways found that "the same model," nominally identical in name and weights, behaves differently depending on which provider serves it, a consequence of infrastructure choices made at the provider level. That finding complicates far more than watermarking: it means benchmark results and evaluation scores can shift based on which access path a request happens to travel through.
Non-determinism occurs even at temperature zero, where output ought to be as repeatable as a system gets. Accuracy varies meaningfully across runs of the same prompt, and some models route tokens internally between specialized subcomponents, a documented architectural factor that adds variance to the output. MoE routing randomizes token distributions at a level that undercuts the consistency a watermark detector needs to build a reliable z-score, since the detector assumes a stable generation process and the routing mechanism doesn't guarantee one.
A related instability appears in the LIDet paper (ICLR 2025), which found that watermarks easily learned by LLMs become unstable with respect to the hash key used during detection. Even if a pipeline never once changes providers or models, it can still produce inconsistent z-scores across separate detection runs on the same underlying content, simply because of this hash-key instability. Put together, these findings mean you can't rely on token-level provenance signals, even if you make no attempt to evade detection and run what you believe is a single, consistent model. The infrastructure disperses fingerprints as an incidental side effect of how models get deployed at scale, not as the result of anyone's evasion strategy. If that dispersal happens by accident under ordinary operating conditions, the only way a content team can know where it actually stands is to design its architecture deliberately and measure the resulting z-scores directly, rather than assume a single-model pipeline guarantees a stable fingerprint.
Regulatory and compliance pressure that lands on content teams operating multi-provider pipelines
Provenance requirements for AI-generated content are already in force in parts of the world, not a future consideration content teams can defer. The EU AI Act has transparency and general-purpose AI provisions enforceable from 2 August 2026, with high-risk system requirements phasing in through 2027 and 2028. The Act requires providers of generative AI systems, general-purpose AI systems included, to mark outputs in a machine-readable format so they can be detected as AI-generated, and it requires deployers to clearly label deepfakes and AI-generated text published on matters of public interest. It also classifies AI systems operating in high-impact sectors, biometrics, critical infrastructure, employment, and law enforcement among them, as high-risk, a classification that can extend to multi-agent systems performing those functions and that triggers requirements for human-in-the-loop oversight, immutable audit trails, and persistent identity management across the agent's lifecycle.
Analysts describe an "illusion of compliance" that takes hold in internal-only environments, where content built during experimentation feels contained right up until it flows into customer-facing output without anyone updating the compliance posture that governed it. Shadow AI workflows, assembled informally by content teams stitching together multiple providers, unmanaged prompts, and no audit trail, are this exact pattern playing out at scale, and they are precisely the kind of pipeline where a watermark's silent collapse would go unnoticed until an audit asks for proof that never existed.
C2PA content credentials offer one path to cryptographic provenance, but they operate at the level of the file, not the token. They work well for an image or a document that stays intact as a single asset moving through a tracked workflow. If content gets copied and pasted across providers without the credential propagating along with it, which is how a multi-provider content pipeline ordinarily works, C2PA offers no protection at that point. It complements token-level watermarking without substituting for it.
The compliance gap is specific: a cross-provider pipeline that has inadvertently destroyed its own z-score has, in the same moment, destroyed the forensic record an audit would ask to see. Measuring the z-score directly and knowing, in advance, what the pipeline's generation architecture does to it is what tells a deliberate, defensible architectural choice apart from an accidental compliance exposure.
Sources
- Published as a conference paper at ICLR 2025 CAN WATERMARKS BE USED
- Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarking
- Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?
- A Watermark for Large Language Models
- A Watermark for Large Language Models
- LLM Watermark Evasion via Bias Inversion
- Towards Safe and Efficient Low-Entropy LLM Watermarking
- Your "Pro" LLM Subscription May Actually Be "Free": Exposing Fingerprint Spoofing Risks in LLM Inference Services


