01 / The mechanism and its boundary
What the technique establishes
DiFR (Divergence From Reference) is a pair of methods for checking that an inference provider ran the model and settings it declared. Both tolerate the small numerical differences that make re-runs disagree. In Token-DiFR, provider and verifier share the random seed for token sampling. The verifier re-runs the sequence and scores how far each claimed token departs from the reference model's choice. Activation-DiFR compares compressed fingerprints of internal activations instead. On models of 8 to 30 billion parameters on A100 and H200 GPUs, both detected 4-bit quantization almost perfectly (AUC above 0.999). Token-DiFR did so within 300 output tokens and Activation-DiFR within 2. The code is public, with a vLLM integration, and another team has built it into a prototype. The verifier needs the model weights and must know the provider's sampling procedure. Its statistical tolerance leaves a covert channel, which an independent study widened in a related exfiltration detector.
- Threat model
- Adversarial prover
- Adversarial evaluation
- Published analysis
- Hardware needed
- None
- Prover cooperation
- Required
- Confidentiality
- Revealing
- Category
- Cryptography & computation
Technical detail and cited results
- The specification φ fixes the architecture and weights, numerical precision, sampling hyperparameters and the seed of the pseudorandom number generator S-0016.
- Token-DiFR works with Gumbel-Max sampling. From the shared seed, both parties derive the same Gumbel noise vector g. The verifier computes post-Gumbel scores z_i = l_i + T·g_i from its reference logits l and temperature T, and takes the token it would have chosen. It scores each claimed token by the logit-margin difference between the two tokens, clipped at a maximum. Scores are pooled per batch. A batch is flagged if the statistic exceeds a threshold set on a calibration set generated under φ on trusted hardware S-0016. Verification is one prefill pass over prompt plus output S-0016. The companion exfiltration paper gives a variant for inverse-probability-transform samplers S-0015.
- Activation-DiFR has both parties generate the same random orthogonal projection P from a shared seed. P is k by D, with k much smaller than D. The provider sends f = P·a for activation vector a, possibly only for every J-th token. The verifier checks the L2 distance to its own projection S-0016. For detecting FP8 KV-cache quantization on H200 GPUs, the authors report an AUC of at least 0.9999 with under 0.5 bytes per token. For Qwen3-30B-A3B on A100 GPUs, an AUC of 0.999 took 7.25 bytes per token, against 32 for TOPLOC S-0016.
- DiFR works with unmodified vLLM, which exposes per-request seeds as a standard sampling parameter S-0016.
- At temperature zero, Token-DiFR needs no seed synchronization. The verifier checks that the provider chose the most likely token at each position. The authors warn that such greedy spot checks are open to selective cheating. A provider could serve the declared model at temperature zero and a degraded one otherwise S-0016.
Claims and scope
A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.
The declared model is the one being served
Checks that outputs are consistent with the declared model, precision and sampling configuration.
Model weights have not left the facility
Used as the estimator in a weight-exfiltration detection scheme (S-0015).
Readiness for a stated use
Assessed use: checking that outputs match the declared model, precision and sampling settings
medium confidence · current · assessed 2026-09-25 · rubric 1.1
This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.
The code is public and the results reproduce on data-centre GPUs, but only its developers rely on it and no one has independently evaluated its consistency check.
- R1 met: the paper states the verification claim, the specification the provider must follow and the trust assumptions S-0016. The assumptions are a trusted reference, calibration on trusted hardware and synchronized seeds. A companion paper embeds the method in a formal security game S-0015.
- R2 met: public MIT-licensed code with a vLLM integration exists S-1005. Results are reproducible on A100 and H200 GPUs with models of 8 to 30 billion parameters S-0016. A separate team built it into a prototype and tested models of up to 120 billion parameters S-1006 S-1007.
- R3 not met: Amodo's prototype is research code, not production software S-1007, and no party is documented relying on DiFR for a verification decision.
- R4 not met: as of September 2026 no independent audit, red-team or peer-reviewed security analysis of DiFR's consistency check has been published. Amodo lists recomputation red-teaming as not started S-1008. Amodo's own tests are independent of the authors but used two fine-tuned models, which by Amodo's account cover only a fraction of possible attacks S-3001. The one independent attack targets a weight-exfiltration detector built on the same Gumbel-margin statistic S-1507. That attack bears on the supporting exfiltration use, assessed under Bounding unexplained information in outputs. It does not bear on the primary use.
Evidence needed for the next level
Reliance by a party other than the developers on DiFR for a verification decision, or a production-grade release.
An independent public security evaluation (audit, red-team or peer-reviewed analysis) against adaptive adversaries.
Limitations, flaws, and blockers
These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.
significant / open / demonstrated attack
Statistical tolerance leaves a covert channel
Statistical schemes can put an upper bound on an adversary's covert bandwidth, but cannot close it S-0020. In the companion exfiltration study, the detector cut exfiltratable information to under 0.5% under benign prompt traffic. It did not cut it to zero S-0015. Its authors called the channel's size under adversarial prompts an open empirical question S-0015. An independent study attacked that detector, which uses the same Gumbel-margin statistic. An adversary who controls the prompts roughly doubled the bits leaked per token. Across six models, this cut the slowdown from 146–254 times under benign prompts to 60–118 times S-1507. The attack widens the exfiltration bound. It does not target DiFR's check that outputs match the declared configuration.
minor / open / open question
Mixed hardware widens the honest baseline
For Qwen3-30B-A3B, pooling honest runs across A100 and H200 GPUs and parallelism setups broadened the honest score distribution. Token-DiFR then failed to separate the two smallest tested changes, a temperature of 1.1 instead of 1.0 and a simulated top-2 sampling bug, at the target false-positive rate, while cross-entropy did. The authors report that matched provider and verifier environments, or pooling that weights rare large deviations, restore detection.
significant / open / open question
Speculative decoding and multi-model sampling not evaluated
The algorithms and experiments cover sampling from a single LLM. Speculative decoding was not evaluated. The authors sketch an extension to one speculative-decoding algorithm, without experiments. They note that other variants would need modified verification and extra metadata.
What still blocks use or stronger assurance
- S-0016
The verifier needs the model weights, so outsiders cannot use the method to verify providers of closed-weights models.
- S-0016S-1006
The verifier must know and match the provider's sampling procedure, and in one prototype a sampling mismatch in a newer vLLM version produced large spurious logit differences.
- S-1008S-1507
No independent red-team of DiFR's consistency check has been published, Amodo rates recomputation red-teaming 'not started', and the one independent attack study targets an exfiltration detector built on the same statistic.
Connections in the research map
Mechanisms implemented
Organizations and developers
The Consortium’s case files
Related editorial reviews use the Consortium’s own descriptive scores and review dates. Their scores are separate from the atlas readiness rubric.
IV-10 / WorkloadThe honest twin serverRead case file ↗Sources and provenance
- S-0016 / Tier B
DiFR: Inference Verification Despite Nondeterminism ↗
A. Karvonen, D. Reuter, R. Rinberg, L. Marks, A. Garriga-Alonso, K. Warr · 2025 · ICML 2026 Workshop on Technical AI Governance Research
Supports: method, specification, experiments, results, comparison with TOPLOC and distributional methods, provider case study, deployment considerations, limitations, speculative-decoding sketch
Locator: abstract; §2-3; §5; §5.1; §6.2; Table 1; §7.2-7.4; Appendix D; Appendix F
Version and catalogue details - S-0015 / Tier B
Verifying LLM Inference to Detect Model Weight Exfiltration ↗
R. Rinberg, A. Karvonen, A. Hoover, D. Reuter, K. Warr · 2025 · arXiv
Supports: companion security game and exfiltration results using Token-DiFR estimators; adversarial prompts an open question
Locator: abstract; contributions; §4; §5.3
Version and catalogue details - S-1005 / Tier B
adamkarvonen/difr (GitHub repository) ↗
A. Karvonen · 2025 · GitHub
Supports: public code, licence, vLLM and API modes
Locator: README
Version and catalogue details - S-1006 / Tier C
Scaling Recomputation Inference Verification ↗
Amodo Design · 2026 · Amodo Design
Supports: independent prototype, scale and vLLM sampling mismatch
Locator: whole note
Version and catalogue details - S-1007 / Tier B
Amodo-Design/Inference-Recomputation-Prototype (GitHub repository) ↗
Amodo Design · 2026 · GitHub
Supports: prototype code, research code not production software; modified copy of the difr library
Locator: README
Version and catalogue details - S-3001 / Tier C
An Inference Verification Prototype — Stage 1 ↗
Amodo Design · 2026 · Amodo Design
Supports: first Amodo prototype reusing the DiFR library; tests against two LoRA fine-tuning attacks on 8×H100; pass-rate gap; short answers; limits of the tests
Locator: setup; results; limitations
Version and catalogue details - S-0020 / Tier B
Bit-Exact AI Inference Verification Without Performance Tradeoffs ↗
N. Cankaya · 2026 · ICML 2026 Workshop on Technical AI Governance Research
Supports: limits of statistical verification
Locator: §1
Version and catalogue details - S-1008 / Tier C
AI 2040 Plan A — Verification SITREP ↗
Amodo Design · 2026 · Amodo Design
Supports: status of red-teaming
Locator: status items
Version and catalogue details - S-1507 / Tier B
Adversarial Entropy Inflation Against Gumbel-Based Inference Verification ↗
N. Kezins · 2026 · arXiv
Supports: independent attack on a weight-exfiltration detector built on the Token-DiFR Gumbel-margin statistic; scope limited to the exfiltration bound
Locator: abstract; introduction
Version and catalogue details
- Source review date
- 2026-09-25
- Drafted by (source map)
- ai
- Review handles (source map)
- codex-review