Model Attribution Research and Its Implications for Content Teams
Watermarking methods fail to meet legal standards for AI attribution.

Model attribution research has reached a point where a single-model content pipeline carries a statistical fingerprint that can be traced back to its source with measurable reliability. The angle that matters for content teams is architectural: how text gets produced now determines how detectable, and how legally defensible, that text turns out to be.
The statistical mechanism behind model attribution
Every token a language model produces comes from a probability distribution shaped by the same sampling bias, which makes single-model output identifiable. Attribution detectors are built to measure exactly this regularity, not to parse meaning or style in the way a human reader would.
The most widely studied mechanism for this is the KGW approach to watermarking. At each decoding step, a hash of the preceding tokens splits the model's vocabulary into two groups, a green list and a red list, and the model's logits for green-list tokens get a fixed boost, often written as δ, before any word is actually chosen. Because the boost is applied before sampling, the bias lives in the probability distribution itself rather than being stitched into the text after the fact. Green tokens appear more often than chance would predict across an entire document generated by that model.
Detecting this is a matter of counting. A detector tallies how many green-list tokens appear in a text and compares that count to what would be expected if the text had no watermark at all, using a statistical test called a one-proportion z-test. The z-score shows how far the observed count strays from what chance would predict. Crossing a threshold causes the detector to flag the text as watermarked. Because the same key and the same context-hashing logic apply to every sentence a given model writes, the green-token excess accumulates steadily across a document. The signal is a property of the architecture that generated the text, not of any individual sentence or word choice. That is why a content team that edits or rewrites individual sentences by hand does not make the signal disappear: the bias sits upstream of the words themselves, embedded in the distribution that selected them before editing ever touches the page.
Fingerprinting works alongside watermarking but answers a different question. Watermarking asks whether any large language model produced a given text. Fingerprinting asks which specific model did. LLMPrint and CoTSRF both build model-unique behavioral signatures, one from prompt-induced token-preference bitstrings, the other from the distribution of a model's chain-of-thought reasoning, and both report high true-positive rates in testing. LLMPrint works in both gray-box and black-box settings, whether or not you have internal access to the model, while CoTSRF is built specifically for black-box conditions where no internal access exists.
None of this comes free. The δ bias that makes detection reliable also degrades the text itself, and that trade-off is a design parameter built into every implementation in the KGW family. Research published in EMNLP 2025 Findings shows that KGW's method of splitting the vocabulary ignores how determinate different parts of speech are: tokens with low entropy, meaning their next word is highly predictable, are where the watermark is hardest to embed without doing visible damage to meaning. That tension, between detection strength and semantic fidelity, is where the next problem begins.
Compliance mandates overstate watermark signal strength and stability
Its reliability falls well short of what regulators and compliance frameworks have assumed when writing watermarking into policy. Tamim et al. (2026) tested three representative methods, KGW, Unigram, and SynthID, against the Daubert admissibility criteria used in U.S. courts and the NIST SP 800-86 digital forensic process standard. To the researchers' knowledge, this was the first study to apply forensic admissibility standards to watermark detection systematically, and the results were not close. All three methods failed before any adversarial attack was even attempted. For KGW specifically, Tamim et al. (2026) found a false-negative rate of 70% on watermarked output that had not been touched by any paraphrasing or evasion attempt.
Paraphrasing compounds this, and it requires no brute-force attack or sophisticated tooling. If you reduce the average conditional probability of sampling a green-list token by even a small margin, detection probability decays exponentially, so evasion follows from the math, not from any clever adversarial technique. Low-entropy token sequences make the failure worse. In domains where the next word is highly predictable, such as domain-specific terminology or structured legal and technical prose, the green-token bias either fails to embed, because the predictable token may sit on the red list and get chosen anyway, or it inflates green-token counts in ordinary human-written text that was never generated by a model. Gu et al. (May 2025) confirm that standard logit-based watermarking, KGW included, breaks down in these low-entropy conditions, where predictable output makes green-token selection unreliable without visibly damaging fluency. This matters directly for content teams working in legal, financial, or technical verticals, where predictable, structured language is the norm.
Attribution ambiguity adds a further layer of instability. A detected watermark signal does not always mean what it appears to mean. It might reflect direct generation by the watermarked model, a trace inherited through knowledge distillation, or deliberate spoofing meant to implicate a model that never touched the text. Pan et al. (Tsinghua, 2025) show that student models trained on a watermarked teacher's outputs inherit detectable watermarks themselves, so a content team using a downstream fine-tuned model could end up carrying provenance signals from a model it never directly queried.
The strongest challenge to this picture comes from fingerprinting, which is more behaviorally stable than token-statistics watermarking and may hold up better for attribution purposes generally. But fingerprinting needs controlled query access and advance knowledge of which model to test against, so it cannot run as a passive scan across content of unknown origin, and that is the scenario compliance teams face most often in practice. The technical foundation, in other words, is fragile in exactly the places where compliance deadlines are now arriving.
What the compliance environment now requires of content workflows
Regulatory obligations around AI content provenance are now active, and most content teams have no documented chain of custody that would hold up against them. The C2PA standard has moved from an emerging initiative into infrastructure in a short span of time. In December 2025, C2PA Specification v2.3 added manifests for unstructured text, so provenance tracking now covers plain-text assets, including large language model output, for the first time. Google has built C2PA Assurance Level 2 directly into Pixel camera hardware. TikTok has implemented mandatory labeling for realistic AI-generated content. These are concrete deployments, not pilot programs.
A structural limitation runs through all of this. C2PA metadata is routinely stripped when content passes through platforms that recompress or reformat files on the way to publication, and that includes most social media networks. The chain of custody breaks precisely at the point of distribution, which is also the point where most content actually reaches its audience.
Google requires a dual-layer provenance standard combining SynthID with C2PA, and OpenAI has converged on the same model, now widely cited as the industry standard for content verification. Google requires AI-generated product images in Merchant Center to retain IPTC DigitalSourceType metadata, and images missing that disclosure can be disapproved. Google does not apply a comparable provenance-based penalty to AI images in organic search results right now, so content teams should not blur that distinction when they weigh their own exposure.
ISO/IEC 42001 sets a traceability bar for AI content teams that is straightforward to state and difficult for most organizations to meet: which prompt, which model, which version, executed at what time, reviewed and approved by whom. Most organizations that run content production across multiple AI tools, agencies, and freelance contributors have none of this documented in any systematic way.
The gap widens further once multi-agent systems enter the picture. Only a minority of enterprises have mature governance frameworks even for single-model deployments, and multi-agent architectures multiply the audit-trail requirement rather than simplify it: the question is no longer just what one model produced, but how a coordinator synthesized the outputs of several agents into a final piece. Microsoft's MDASH, a multi-model agentic scanning harness, offers a working illustration of what a properly instrumented multi-agent system looks like in practice. Parallel specialist agents, structured as auditors and debaters, run through a pipeline designed to discover, debate, and prove exploitable issues, with the coordination itself logged and reviewable. That pattern, instrumented coordination across multiple agents, is rare in enterprise content operations today, even though it is precisely what the compliance environment now demands.
The case against single-model pipelines
A single-model pipeline concentrates the exact statistical fingerprint that attribution systems are built to find. Every token comes from the same probability distribution, shaped by the same green-list bias, and the cumulative z-score across a long document is a direct function of how many tokens that one model generated. There is no amount of editing that fixes this, because the fingerprint is a structural property of the generation architecture, not a surface feature of the prose. Mixed-source detection methods such as GCD and AOL, both developed within the KGW line of watermark research, can localize watermarked spans within a larger document, so even a mostly human-edited piece with a few AI-generated paragraphs can be flagged and attributed to its source.
Dispersing authorship across multiple models at the token level is the structurally sound response to this problem, and it is not a novel proposal. It reflects how enterprise AI systems are already built. Dynamic routers and tree-based logic already direct individual requests to whichever model suits them best, and mixture-of-experts design has become the standard architecture at scale, with total parameter counts reaching hundreds of billions and some frontier models exceeding one trillion parameters, even as the number of parameters actually active on any single forward pass stays manageable. If you build a content workflow on this logic, one model might generate an initial draft, a second refine tone and nuance, and a third handle factual review, so each model leaves a distinct distributional signature on the finished document instead of one model's fingerprint running through the whole thing.
Token-level blending across providers is what actually disperses the green-list signal. When tokens are drawn from multiple models operating under different keys, or from models with no watermarking scheme at all, the green-token count across the document stops accumulating consistently enough to cross the detection threshold. This is genuine multi-model authorship, diversity built in at the level of individual token choices, and it is fundamentally different from paraphrasing or post-hoc editing, approaches the research already shows fail to reliably remove the underlying signal.
Attribution results vary by content domain, which complicates any single blanket strategy. On the Human-AI Parallel Corpus, tested across six domains, style-embedding methods and LLM-judge methods perform differently depending on content type: fiction and academic prose favor the LLM-judge approach, while spoken and scripted dialogue favor embedding-based methods. Attribution is not a single-number problem, and content teams working in different verticals face meaningfully different exposure profiles as a result.
The strongest objection to a multi-model approach is practical: coordinating several models is harder to govern and audit than relying on one, and that concern is legitimate on its own terms. But the governance gap documented in multi-agent systems is a process problem, not an argument against the architecture itself. The Microsoft MDASH pattern demonstrates that instrumented multi-agent pipelines, with coordinator-level audit trails, are buildable today. The alternative, a single-model pipeline that stays structurally identifiable by design, still leaves governance unresolved. It just trades one risk for another: it concentrates detectability instead of spreading it across a documented, auditable process.
What content teams should assess and instrument now
The decisions content teams face about AI writing architecture are now technical and legal decisions, not editorial preferences left to a style guide. First, audit what the current pipeline's statistical signature actually looks like.
That audit begins by mapping single-model exposure across existing workflows. Every step where one model generates a continuous span of tokens is a step that accumulates z-score, and the longer that span runs from a single model, the more detectable the resulting document becomes. Mixed-source detection methods already published in the research can localize these spans inside a larger document, so content teams should assume that a competent detector can do what the published literature says it can do.
A documented chain of custody is the next layer, and it needs to satisfy the traceability requirements set out in ISO/IEC 42001: which prompt was used, which model and version generated the output, when it ran, and who reviewed and approved it. That documentation is the minimum layer separating a defensible workflow from an auditable gap, and coordinator logs in a multi-agent pipeline produce it as a natural byproduct of running the system.
Any claim about removing or reducing watermark detectability needs to be evaluated on quantitative grounds, specifically z-score outcomes measured against real KGW-family detectors, not against paraphrase-diversity scores or perplexity metrics that measure something else entirely and that a court or regulator would not accept as evidence of compliance. The forensic readiness framework built by Tamim et al. (2026), the FRS system for evaluating watermark forensic admissibility under Daubert and NIST SP 800-86 standards, gives content teams a systematic way to stress-test these vendor claims. The same criteria that exposed the gap between regulatory assumption and empirical watermark performance can be turned directly on any tool a content team is considering, and every claim should be held to that standard going forward.
Sources
- Watermarking with Low-Entropy POS-Guided Token ...
- Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?
- AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation
- Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarking
- A Survey of Text Watermarking in the Era of Large Language Models
- Downstream Trade-offs of a Family of Text Watermarks
- Pioneering Efficient Detection of Watermarked Segments in ...
- Towards Possibilities & Impossibilities of AI-generated Text Detection: A Survey


