I-0002 / Cryptography & computation

DiFR (Divergence From Reference)

DiFR checks that an inference provider ran its declared model by comparing output tokens or activations with a trusted re-run using the same random seed.

R2 DemonstratedSource reviewed 2026-09-25Provider-reported evidence

01 / The mechanism and its boundary

What the technique establishes

DiFR (Divergence From Reference) is a pair of methods for checking that an inference provider ran the model and settings it declared. Both tolerate the small numerical differences that make re-runs disagree. In Token-DiFR, provider and verifier share the random seed for token sampling. The verifier re-runs the sequence and scores how far each claimed token departs from the reference model's choice. Activation-DiFR compares compressed fingerprints of internal activations instead. On models of 8 to 30 billion parameters on A100 and H200 GPUs, both detected 4-bit quantization almost perfectly (AUC above 0.999). Token-DiFR did so within 300 output tokens and Activation-DiFR within 2. The code is public, with a vLLM integration, and another team has built it into a prototype. The verifier needs the model weights and must know the provider's sampling procedure. Its statistical tolerance leaves a covert channel, which an independent study widened in a related exfiltration detector.

Threat model
Adversarial prover
Adversarial evaluation
Published analysis
Hardware needed
None
Prover cooperation
Required
Confidentiality
Revealing
Category
Cryptography & computation
Technical detail and cited results
  • The specification φ fixes the architecture and weights, numerical precision, sampling hyperparameters and the seed of the pseudorandom number generator S-0016.
  • Token-DiFR works with Gumbel-Max sampling. From the shared seed, both parties derive the same Gumbel noise vector g. The verifier computes post-Gumbel scores z_i = l_i + T·g_i from its reference logits l and temperature T, and takes the token it would have chosen. It scores each claimed token by the logit-margin difference between the two tokens, clipped at a maximum. Scores are pooled per batch. A batch is flagged if the statistic exceeds a threshold set on a calibration set generated under φ on trusted hardware S-0016. Verification is one prefill pass over prompt plus output S-0016. The companion exfiltration paper gives a variant for inverse-probability-transform samplers S-0015.
  • Activation-DiFR has both parties generate the same random orthogonal projection P from a shared seed. P is k by D, with k much smaller than D. The provider sends f = P·a for activation vector a, possibly only for every J-th token. The verifier checks the L2 distance to its own projection S-0016. For detecting FP8 KV-cache quantization on H200 GPUs, the authors report an AUC of at least 0.9999 with under 0.5 bytes per token. For Qwen3-30B-A3B on A100 GPUs, an AUC of 0.999 took 7.25 bytes per token, against 32 for TOPLOC S-0016.
  • DiFR works with unmodified vLLM, which exposes per-request seeds as a standard sampling parameter S-0016.
  • At temperature zero, Token-DiFR needs no seed synchronization. The verifier checks that the provider chose the most likely token at each position. The authors warn that such greedy spot checks are open to selective cheating. A provider could serve the declared model at temperature zero and a degraded one otherwise S-0016.

Claims and scope

A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.

Readiness for a stated use

R2 Demonstrated

Assessed use: checking that outputs match the declared model, precision and sampling settings

medium confidence · current · assessed 2026-09-25 · rubric 1.1

This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.

The code is public and the results reproduce on data-centre GPUs, but only its developers rely on it and no one has independently evaluated its consistency check.

  • R1 met: the paper states the verification claim, the specification the provider must follow and the trust assumptions S-0016. The assumptions are a trusted reference, calibration on trusted hardware and synchronized seeds. A companion paper embeds the method in a formal security game S-0015.
  • R2 met: public MIT-licensed code with a vLLM integration exists S-1005. Results are reproducible on A100 and H200 GPUs with models of 8 to 30 billion parameters S-0016. A separate team built it into a prototype and tested models of up to 120 billion parameters S-1006 S-1007.
  • R3 not met: Amodo's prototype is research code, not production software S-1007, and no party is documented relying on DiFR for a verification decision.
  • R4 not met: as of September 2026 no independent audit, red-team or peer-reviewed security analysis of DiFR's consistency check has been published. Amodo lists recomputation red-teaming as not started S-1008. Amodo's own tests are independent of the authors but used two fine-tuned models, which by Amodo's account cover only a fraction of possible attacks S-3001. The one independent attack targets a weight-exfiltration detector built on the same Gumbel-margin statistic S-1507. That attack bears on the supporting exfiltration use, assessed under Bounding unexplained information in outputs. It does not bear on the primary use.

Evidence needed for the next level

  • Reliance by a party other than the developers on DiFR for a verification decision, or a production-grade release.

  • An independent public security evaluation (audit, red-team or peer-reviewed analysis) against adaptive adversaries.

Limitations, flaws, and blockers

These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.

significant / open / demonstrated attack

Statistical tolerance leaves a covert channel

Statistical schemes can put an upper bound on an adversary's covert bandwidth, but cannot close it S-0020. In the companion exfiltration study, the detector cut exfiltratable information to under 0.5% under benign prompt traffic. It did not cut it to zero S-0015. Its authors called the channel's size under adversarial prompts an open empirical question S-0015. An independent study attacked that detector, which uses the same Gumbel-margin statistic. An adversary who controls the prompts roughly doubled the bits leaked per token. Across six models, this cut the slowdown from 146–254 times under benign prompts to 60–118 times S-1507. The attack widens the exfiltration bound. It does not target DiFR's check that outputs match the declared configuration.

S-0020S-0015S-1507

minor / open / open question

Mixed hardware widens the honest baseline

For Qwen3-30B-A3B, pooling honest runs across A100 and H200 GPUs and parallelism setups broadened the honest score distribution. Token-DiFR then failed to separate the two smallest tested changes, a temperature of 1.1 instead of 1.0 and a simulated top-2 sampling bug, at the target false-positive rate, while cross-entropy did. The authors report that matched provider and verifier environments, or pooling that weights rare large deviations, restore detection.

S-0016

significant / open / open question

Speculative decoding and multi-model sampling not evaluated

The algorithms and experiments cover sampling from a single LLM. Speculative decoding was not evaluated. The authors sketch an extension to one speculative-decoding algorithm, without experiments. They note that other variants would need modified verification and extra metadata.

S-0016

What still blocks use or stronger assurance

  1. The verifier needs the model weights, so outsiders cannot use the method to verify providers of closed-weights models.

    S-0016
  2. The verifier must know and match the provider's sampling procedure, and in one prototype a sampling mismatch in a newer vLLM version produced large spurious logit differences.

    S-0016S-1006
  3. No independent red-team of DiFR's consistency check has been published, Amodo rates recomputation red-teaming 'not started', and the one independent attack study targets an exfiltration detector built on the same statistic.

    S-1008S-1507

Connections in the research map

Mechanisms implemented

Concepts used

Organizations and developers

The Consortium’s case files

Related editorial reviews use the Consortium’s own descriptive scores and review dates. Their scores are separate from the atlas readiness rubric.

IV-10 / WorkloadThe honest twin serverRead case file ↗

Sources and provenance

  1. S-0016 / Tier B

    DiFR: Inference Verification Despite Nondeterminism ↗

    A. Karvonen, D. Reuter, R. Rinberg, L. Marks, A. Garriga-Alonso, K. Warr · 2025 · ICML 2026 Workshop on Technical AI Governance Research

    Supports: method, specification, experiments, results, comparison with TOPLOC and distributional methods, provider case study, deployment considerations, limitations, speculative-decoding sketch

    Locator: abstract; §2-3; §5; §5.1; §6.2; Table 1; §7.2-7.4; Appendix D; Appendix F

    Version and catalogue details
  2. S-0015 / Tier B

    Verifying LLM Inference to Detect Model Weight Exfiltration ↗

    R. Rinberg, A. Karvonen, A. Hoover, D. Reuter, K. Warr · 2025 · arXiv

    Supports: companion security game and exfiltration results using Token-DiFR estimators; adversarial prompts an open question

    Locator: abstract; contributions; §4; §5.3

    Version and catalogue details
  3. S-1005 / Tier B

    adamkarvonen/difr (GitHub repository) ↗

    A. Karvonen · 2025 · GitHub

    Supports: public code, licence, vLLM and API modes

    Locator: README

    Version and catalogue details
  4. S-1006 / Tier C

    Scaling Recomputation Inference Verification ↗

    Amodo Design · 2026 · Amodo Design

    Supports: independent prototype, scale and vLLM sampling mismatch

    Locator: whole note

    Version and catalogue details
  5. S-1007 / Tier B

    Amodo-Design/Inference-Recomputation-Prototype (GitHub repository) ↗

    Amodo Design · 2026 · GitHub

    Supports: prototype code, research code not production software; modified copy of the difr library

    Locator: README

    Version and catalogue details
  6. S-3001 / Tier C

    An Inference Verification Prototype — Stage 1 ↗

    Amodo Design · 2026 · Amodo Design

    Supports: first Amodo prototype reusing the DiFR library; tests against two LoRA fine-tuning attacks on 8×H100; pass-rate gap; short answers; limits of the tests

    Locator: setup; results; limitations

    Version and catalogue details
  7. S-0020 / Tier B

    Bit-Exact AI Inference Verification Without Performance Tradeoffs ↗

    N. Cankaya · 2026 · ICML 2026 Workshop on Technical AI Governance Research

    Supports: limits of statistical verification

    Locator: §1

    Version and catalogue details
  8. S-1008 / Tier C

    AI 2040 Plan A — Verification SITREP ↗

    Amodo Design · 2026 · Amodo Design

    Supports: status of red-teaming

    Locator: status items

    Version and catalogue details
  9. S-1507 / Tier B

    Adversarial Entropy Inflation Against Gumbel-Based Inference Verification ↗

    N. Kezins · 2026 · arXiv

    Supports: independent attack on a weight-exfiltration detector built on the Token-DiFR Gumbel-margin statistic; scope limited to the exfiltration bound

    Locator: abstract; introduction

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review