APIs, integration & security — in depth

How Token Distribution Differs Between GPT-4 and Claude

Different tokenizers create divergent probability distributions that undermine detection methods.

Senior Editor, AI Compliance · · 9 min read
Cover illustration for “How Token Distribution Differs Between GPT-4 and Claude”
LLM Fingerprinting · October 6, 2026 · 9 min read · 2,101 words

Most discussions of GPT-4 and Claude start with personality. GPT-4 is called more clinical, Claude more conversational, and the comparison usually ends at tone, verbosity, or benchmark scores like the kind trade publications run every time a new version ships. These comparisons describe real effects, but they describe them at the wrong layer. The personality of a model's output is downstream of something that happens before a single word is chosen: the text has to be broken into tokens, and GPT-4 and Claude do not break it the same way. That divergence in tokenization is the mechanical origin of nearly everything that follows, from how each model prices a request to how a forensic classifier tells the two apart. Any argument about detection, attribution, or watermarking that skips this layer is building on a surface observation.

A token is the unit a language model actually computes over. Every probability the model assigns, every statistical signature a detector later tries to read, is calculated at the token level, not at the level of words or sentences a human reader perceives. Because GPT-4 and Claude use different tokenizers, the same sentence fed into each model produces a different sequence of token IDs. That means the probability distributions each model generates, the green-list assignments a watermark might apply, and the stylometric signals a forensic tool might extract are all computed over different vocabularies from the very first step. Nothing downstream can be understood correctly if this starting divergence is ignored.

How GPT-4 and Claude tokenize the same text differently

The practical meaning of "different tokenizers" is that identical English text segments differently depending on which model is doing the segmenting. The boundaries where one token ends and the next begins do not line up between GPT-4 and Claude, and the numerical IDs assigned to those tokens come from entirely separate vocabularies. This is a difference in both where the cuts fall and what the resulting pieces are labeled.

Verifying this directly for Claude is harder than it sounds, because Anthropic has not published an accurate, reproducible tokenizer for Claude 3 or later models that developers can run locally. Third-party estimates of Claude's token counts are built from character-to-token ratios rather than from a reproduction of the actual segmentation logic. The commonly cited heuristic, roughly 3.5 characters per token for English text on older Claude Sonnet and Opus models before version 4.7, is a rough planning estimate rather than a guarantee, and it diverges from Anthropic's own internal count as a conversation approaches its context limit.

For most everyday text, Claude's token count and GPT's token count land in a similar range, close enough that a content team estimating cost or context budget can use either as a rough guide. But a count that looks close conceals a deeper mismatch: two sequences with similar lengths are still composed of different tokens, carrying different identities in each model's respective vocabulary. A sentence that costs a certain number of tokens under one tokenizer may cost a different number under the other, and even where the totals happen to match, the individual tokens themselves do not correspond to each other one for one.

That overhead compounds in ways invisible to the end user. Claude processes system instructions on every request that the user never sees, and those instructions carry a real token cost each time the conversation is replayed. OpenAI maintains analogous hidden system content that it does not publish either. Because each new message in a conversation is processed by replaying the full history plus whatever system prompts apply, even a short new message carries the accumulated token weight of everything that came before it in the session. The upshot of all of this is simple to state and significant in its consequences: GPT-4 and Claude do not just count differently, they operate over different token ID sequences drawn from different vocabulary spaces. That gap in vocabulary space is where the next layer of the problem, watermark detection, begins to break down.

KGW Watermarking and Token Boundaries

The most widely used framework for detecting AI-generated text, the KGW scheme, works entirely at the token level. That design choice means a tokenization divergence between two models is not a cosmetic difference, it is a structural fault line running directly through the detection apparatus.

KGW watermarking works by splitting a model's vocabulary into a green list and a red list for each position in a generated sequence, then nudging the model's output probabilities to favor green-list tokens. The split itself is not arbitrary: for each token the model generates, the algorithm hashes the preceding context and uses that hash to deterministically decide which tokens count as green and which count as red at that position. The same context will always produce the same green/red split for a given model's vocabulary. Detection then works backward from a finished text by counting how many of its tokens landed on the green list and computing a z-score, which measures how far that green-token count deviates from what plain chance would predict. A high z-score is the statistical signature of a watermark having been applied.

The entire scheme depends on the green-list partition being meaningful for the vocabulary it was computed against. GPT-4 and Claude draw from different token vocabularies, so a green-list partition built for one model's vocabulary carries no coherent meaning when applied to text generated by the other. When the actual green-token ratio in a text diverges from the ratio the detector expects for a given hash key and vocabulary, z-score detection with a fixed threshold stops working reliably, even where a genuine watermark was applied correctly under its own model's scheme. This failure is the predictable consequence of applying a method built on a single-model assumption, suited to a single model generating text under its own vocabulary, to an environment where more than one vocabulary is in play.

Increasing the logit bias applied to green-list tokens makes the watermark easier to detect, but it degrades the quality of the generated text. That trade-off is calibrated against a single model's probability space. Once a different vocabulary enters the picture, that calibration no longer applies in either direction, and the detection threshold built around one model's quality-detectability balance has nothing reliable to measure in a cross-model context.

What different token distributions look like as statistical fingerprints

The vocabulary split between GPT-4 and Claude does more than disrupt a single detection scheme built around watermarking. It leaves behind statistical fingerprints that attribution researchers have learned to extract and read independently of any watermark.

Word-level distributional patterns let embedding-based classifiers distinguish major model families, including GPT, Claude, and Gemini, and these classifiers keep working even after a text has been rewritten or translated. The reason this robustness holds is that the signal lives in the shape of the distribution, the statistical tendencies across many tokens, rather than in any particular word choice that a paraphrase could simply remove. These signals are sensitive to the prompt that generated the text and to the decoding settings used at generation time. Researchers treat model attribution as probabilistic evidence about origin.

Even the distribution of a single output token carries enough information to serve as a probabilistic behavioral fingerprint. That fingerprint is strong enough to give outside observers meaningful evidence about whether an API endpoint is actually serving the model it claims to be running, though the method carries a known equal error rate and a known vulnerability to deliberate spoofing by an adversary who understands the technique.

Several distinct approaches have grown out of this line of work, and each one interacts with the vocabulary split between models in its own way. Black-box behavioral fingerprinting works directly from single-token output distributions and requires nothing more than ordinary API access, with no need to see inside the model itself. Chain-of-thought fingerprinting instead treats the reasoning trace a model produces as a stealthier and more robust signal than the final answer alone, since the trace carries distributional information the finished answer may not retain. Gradient-based fingerprinting reaches further, exploiting signals in the model's weight space, a method unavailable to anyone working only through an API but accessible to parties with direct model access. Behavioral and refusal vectors offer another angle, reading how a model declines or hedges a request as a provenance clue in its own right. Natural distributional fingerprints round out the picture: identifiable signatures turn up even among models trained on the same dataset, a byproduct of randomness in the optimization process, appearing even among models trained on identical data, so that models sharing identical training data can often still be told apart above chance, though the effect grows weaker as the models in question grow larger. What unites all of these approaches is that the underlying signal is architectural. It is rooted in how each model's vocabulary and token distribution were built.

Mixing GPT-4 and Claude Outputs Breaks Watermarks and Attribution

Content teams routinely route tasks across models as a matter of workflow: GPT-4 handling structure and logical scaffolding, Claude handling prose and voice, and the two outputs stitched into a single finished document. That kind of routing is a rational response to each model's relative strengths, and it happens constantly in production content pipelines.

It also produces a document whose token distribution belongs to neither model. When the outputs of GPT-4 and Claude are composed into one document, whether by deliberate editorial routing or by an agentic pipeline switching models mid-task, the resulting text draws tokens from two separate vocabularies at once. A watermark detector calibrated to a single model's green-list assignments has no coherent baseline to measure against in that composite text, and the watermark signal tends to cancel out or dilute below the threshold the detector needs to register a positive result.

This dilution comes from the same structural mismatch described above, now applied to a concrete production scenario. KGW-style detection assumes one vocabulary and one green-list partition governing the entire text under examination. A document assembled from two models' outputs has a green-token ratio that fits neither model's expected distribution, because no single partition was ever applied consistently across the whole piece. The single-model assumption built into KGW-family detectors fails precisely in the multi-model case, which is exactly the case that routine content workflows now produce as a matter of course.

Attribution suffers an analogous breakdown. Distributional classifiers can often identify which model family dominates a piece of text, but a genuinely mixed document produces signals that no current method resolves with certainty. The practical consequence for compliance and provenance tracking is direct: infrastructure built around single-model pipelines, including cryptographic content credentials meant to establish a clean chain of custody, loses that chain the moment more than one model's vocabulary contributes to the finished text. There is no single token space left to compute a consistent provenance record over, because the document was never generated inside one.

Architectural diversity, not paraphrasing, as the structurally correct response

If watermark detectability depends on consistent green-list assignments inside a single model's vocabulary, then the only response that addresses the mechanism directly is to generate content across genuinely distinct vocabularies, not to rephrase text that still originated from one model's token space.

Paraphrasing and other post-hoc editing passes are already recognized as a primary robustness challenge for token-level watermarking, because regenerating or substituting tokens through a different process can meaningfully disrupt the green/red assignments a watermark relies on and degrade its statistical signature. But paraphrasing within a single model's output does not change the fact that every token in that output was drawn from one vocabulary, weighted by one model's probability distribution. Editing the surface text after the fact does nothing to the vocabulary that generated it or to the green-list partition that was applied during that generation.

Dispersing generation across providers whose tokenizers diverge at the vocabulary level works differently, and more fundamentally. A composite text built from genuinely separate token vocabularies produces green-token ratios that are incoherent with respect to any single model's watermark schema, because there is no single hash partition that applies across the whole piece. This is the same mechanism responsible for the attribution and provenance problems described above, now understood as a structural property. The vocabulary divergence between GPT-4 and Claude that makes clean attribution difficult is the same divergence that makes genuine cross-model generation resistant to single-model watermark detection by construction. Understanding tokenization as the root layer of model behavior is what makes that connection visible, and it is the only vantage point from which claims about detection, attribution, or watermark compliance can be evaluated with any precision.

Sources

  1. Published as a conference paper at ICLR 2025 CAN WATERMARKS BE USED

More in LLM Fingerprinting