Skip to content

One Token Is Enough: Fingerprinting and Verifying Large Language Models from Single-Token Output Distributions

Jul 2026 · arXiv.org · Vol abs/2607.10252 · 8 citations · ⚡ 1 influential · 23 references
Computer Science

TL;DR

A behavioral fingerprint of an LLM is defined as the empirical distribution of its answers to trivial one-word prompts collected across four languages at a cost of one output token per query, and it is found that these distributions are highly non-uniform and model-specific.

Abstract

Large language models (LLMs) are increasingly consumed through opaque serving chains - API aggregators, resellers, and inference providers - in which the client has no technical means to confirm that the model answering is the model advertised, and recent audits show that a substantial fraction of commercial endpoints deviate from the vendor's reference weights. Existing identification techniques require long generated texts, token-level log-probabilities, adversarially crafted prompts, or the model owner's cooperation. We show that far weaker evidence suffices. We define a behavioral fingerprint of an LLM as the empirical distribution of its answers to trivial one-word prompts -"name a random number between 1 and 100"- collected across four languages at a cost of one output token per query. Measuring 165 models served via a large commercial aggregator (OpenRouter), we find that (i) these distributions are highly non-uniform (median cell entropy 1.0 bit) and model-specific: split halves of the same model's samples lie an order of magnitude closer than samples of different models; (ii) Jensen-Shannon divergence between fingerprints recovers model lineage, assigning a model to its documented family with 59.5% leave-one-out accuracy against an 18.4% chance rate; and (iii) a biometric-style verification protocol achieves a 7.3% equal error rate with the full 40-cell battery, and below 11% with eight probe cells - roughly a hundred single-token queries per audit. We further report ecosystem anomalies, including a proprietary-branded flagship endpoint distributionally indistinguishable from an open-weight Qwen model. The protocol, prompts, raw data, and analysis code are released for reproduction and operational use.

View source

Similar papers

#natural language process... Preprint Aug 2026

Token Counts Are Not Model Lineage: A Frozen-Threshold Holdout Study of Black-Box LLM API Fingerprinting

A validity-gated result contract is introduced that distinguishes an observed dissimilarity from an uninformative measurement caused by missing usage data, rate limits, or endpoint policy, and validates token-count consistency as a fingerprint of a shared tokenization stack, but rejects its use as a standalone necessar...

Bolin Chen · 0 citations
Preprint Sep 2026

Beyond QA Matching: Perturbation-Response Fingerprinting via Probability Distributions for Large Language Models

Large language models are often instruction-tuned, specialized, quantized, or otherwise transformed, making fine-grained provenance difficult. In this paper, we introduce BReF, a training-free fingerprint that compares how probability distributions over four answer-option labels A/B/C/D move under controlled textual pe...

Ji-Chao Zeng, Yan-Li Chen, Han-Zhou Wu · 0 citations
#natural language process... Preprint Sep 2026

Where a Model Sends Its Own Repeated Token

Black-box model identification works by scoring a model's response to natural-language prompts. One line of work feeds models a degenerate input -- their own token, repeated -- to find a failure mode rather than an identity. We take that input and ask where the model goes when it does not. For each token t, read argmax...

Nicolás Vera Zúñiga · 0 citations
Preprint Sep 2026

Practical Secrets Extraction against Black-box LLMs

Large language models (LLMs) increasingly power autonomous coding agents such as Codex and Claude Code, yet their training corpora may contain confidential credentials exposed in public repositories or collected from private development artifacts, creating risks of memorization and subsequent leakage. Existing extracti...

Shi-Qian Zhao, Si-Wei Jiang, Xin-Feng Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Bits to Beliefs: Recoverable Semantic Fingerprints for Black-Box Verification of Large Language Models

Open-weight large language models (LLMs) can be copied, modified, and redeployed behind black-box APIs, making post-release ownership verification difficult. Existing black-box fingerprints often rely on secret query-key pairs that reproduce predefined responses, and can therefore be easily disrupted by fine-tuning, pr...

Jia-Xin Hong, Yu-Xin Peng, Hong-Yao Yu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Controlled Decoding Attacks on Black-Box LLMs

This work introduces \method{}, a framework for jailbreaking through text-only continuation interfaces that permit repeated sampling and assistant-prefix continuation, and achieves the highest mean score most comparisons against baselines.

Jesson Wang, Shawn Li, Wei Yang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.