01 / The mechanism and its boundary
What the technique establishes
Sampled inference recomputation is a check that an AI provider's recorded outputs came from the model and settings it declared. The provider logs each query and response. A verifier's server re-runs a random sample on a trusted copy of the model. It flags outputs that differ by more than benign numerical noise allows. Checking costs less than generating because the verifier processes each sequence in one pass. Two schemes, TOPLOC and DiFR, have public code. DiFR has reproducible results on A100 and H200 GPUs with models of 8 to 30 billion parameters. TOPLOC's developer has used it to accept or reject work from untrusted computers. No independent red-team of the check has been published. For closed-weights models, the verifier's copy must sit in a trusted, confidential environment. Unrecorded work is never checked. The noise tolerance also leaves a covert channel. An adversary who controls the prompts can widen it.
- Threat model
- Adversarial prover
- Adversarial evaluation
- Published analysis
- Hardware needed
- None
- Prover cooperation
- Required
- Confidentiality
- Partial
- Category
- Cryptography & computation
Technical detail and cited results
The reference protocol has five steps S-0015:
- The serving stack writes every request and response to an append-only log, with the model release, code version and sampling seed S-0015.
- An isolated verification server draws a private random sample of logged pairs. Its isolation can range from access controls and containers to a full air gap S-0015.
- The server runs one prefill pass over the prompt and claimed output with the trusted model and code. This gives the next-token distribution at every position S-0015 S-0016.
- The server scores how far the claimed output diverges from the reference. Token-DiFR uses the clipped Gumbel-Max logit margin under a shared seed. Activation-DiFR uses the distance between random projections of activations S-0016. TOPLOC counts exponent mismatches and measures mantissa differences in the top-128 last-layer activations S-1000.
- Scores beyond a threshold calibrated on trusted hardware are escalated for review S-0015 S-0016.
Verification is cheap because it needs only the prefill pass. Decoding typically runs at 3–5 times lower hardware utilization than prefill S-0015. Amodo's prototype measured a 2–8 times performance advantage for the verifier with models of 1.5 to 120 billion parameters on H100 and H200 GPUs S-1006.
The number of samples sets the chance of detection. For a false-output rate p, the chance of catching at least one false output in k independent samples is 1 − (1 − p)^k S-0020. With 3,000 samples, a 0.1% rate is caught with 95.03% probability S-0018.
Recomputation can also match bit for bit, which makes the check pass/fail. This exact variant needs a known hardware model, deployed weights, parallelism topology, software versions and per-pass batch size, and no atomic operations S-0018 S-0020.
Claims and scope
A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.
The declared model is the one being served
Checks that sampled recorded outputs are consistent with the declared model, precision and sampling settings.
Model weights have not left the facility
Bounds how much information can be hidden steganographically in checked outputs. It is not a stand-alone defence against weight exfiltration (S-0015).
This compute runs inference, not training
Proposed as the correctness check in inference-only retrofit plans. Completeness needs other mechanisms (S-0067, S-0017).
Readiness for a stated use
Assessed use: checking untrusted workers' activations against the declared model, prompt and precision
low confidence · current · assessed 2026-10-08 · rubric 1.1
This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.
TOPLOC has been used in production by its developer, but no independent security audit or red-team of TOPLOC has been published.
- R1 met: the design, the claim it verifies and its trust assumptions are published. They include a formal security game for steganographic weight exfiltration S-0015.
- R2 met through DiFR and TOPLOC. DiFR has MIT-licensed code with a vLLM integration S-1005. It has reproducible results on A100 and H200 GPUs with models of 8 to 30 billion parameters S-0016. TOPLOC is an installable package S-1001 with peer-reviewed results S-1000. Rinberg et al. evaluate the method against a stated adversary who controls the inference server S-0015.
- R3 met through TOPLOC: it is publicly available S-1001, and Prime Intellect, its developer, reports using it in two 2025 runs to accept or reject work from untrusted nodes S-1003 S-3000. No other party is documented relying on sampled recomputation for a verification decision. This checks forward-pass consistency. Token-level checks of sampling settings through DiFR remain at R2 S-0016.
- R4 not met: as of September 2026 no independent audit, red-team or peer-reviewed security analysis of the consistency check has been published. Amodo lists recomputation red-teaming as not started S-1008. Amodo's tests of Token-DiFR are independent of its authors but used two fine-tuned models, which by Amodo's account cover only a fraction of possible attacks S-3001. The one independent attack study targets the weight-exfiltration bound built on this mechanism S-1507. That bound is a supporting use, assessed under Bounding unexplained information in outputs.
Confidence is low: R3 rests on the developer's own reports of TOPLOC use, the latest from July 2025 S-3000.
Evidence needed for the next level
An independent audit, red-team or peer-reviewed security analysis of a recomputation consistency check against an adaptive adversary.
Limitations, flaws, and blockers
These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.
significant / open / demonstrated attack
Tolerance for numerical noise leaves a covert channel
Schemes that accept approximate matches can put an upper bound on an adversary's covert bandwidth, but they cannot close the channel S-0020. The weight-exfiltration detector cut exfiltratable information to under 0.5%, not to zero, on a 30-billion-parameter mixture-of-experts model under benign prompt traffic S-0015. Its authors called the channel's size under adversarial prompts an open empirical question S-0015. An independent study showed that an adversary who controls the prompts roughly doubles the bits leaked per token. Across six models, that cut the slowdown from 146–254 times under benign prompts to 60–118 times S-1507. The attack widens the exfiltration bound. It does not target the check that outputs match the declared model S-1507.
significant / open / theoretical argument
Only recorded traffic is checked
Recomputation checks that recorded, declared workloads are correct. It cannot show that the record is complete. The published schemes do not cover hidden workloads run on the same compute, or substituted work S-0017. Rinberg et al. say their exfiltration-detection scheme cannot stand alone S-0015.
significant / open / theoretical argument
Some inference optimizations are not covered
TOPLOC's authors state that it cannot detect speculative decoding in which a cheaper model does the decoding. They did not test whether it distinguishes types of key-value (KV) cache compression S-1000. DiFR was evaluated only on sampling from a single model. Its authors sketch an extension to one speculative-decoding algorithm but do not test it S-0016.
minor / open / open question
Mixed hardware widens the honest baseline
When honest reference runs span different GPU types, the spread of benign scores grows. In DiFR's tests on Qwen3-30B-A3B, pooling A100 and H200 runs left Token-DiFR unable to separate the two smallest tested changes, a temperature of 1.1 instead of 1.0 and a simulated top-2 sampling bug, at the target false-positive rate, while cross-entropy separated them. Matched provider and verifier environments, or pooling that weights rare large deviations, restored detection S-0016.
What still blocks use or stronger assurance
In tap-based retrofit designs, recording all inference traffic needs network taps and recomputation servers that can ingest it, in the worst case one recomputation-server network interface per inference front-end interface.
Dependency: Network taps and certifiers
S-1008S-0067- S-1008S-0015
In retrofit designs, the recomputation server must sit inside the prover's data centre, possibly under the prover's physical control, and still be protected from a compromised provider, which Amodo rates 'not on track'.
- S-1008S-1507
No independent red-team of a recomputation consistency check has been published (the one independent attack study targets the weight-exfiltration bound), and Amodo rates recomputation red-teaming 'not started'.
- S-0016S-1006
Tolerance-based checks need calibration on trusted hardware and exact knowledge of the provider's sampling procedure, and in one prototype a sampling-implementation mismatch produced large spurious differences.
- S-0016S-0017S-0018
The verifier needs the model weights, so checking a closed-weights model requires a trusted, confidential recomputation environment, which the retrofit designs place inside the prover's facility.
Connections in the research map
Complementary techniques
Alternative approaches
Concepts used
Organizations and developers
Implementations
The Consortium’s case files
Related editorial reviews use the Consortium’s own descriptive scores and review dates. Their scores are separate from the atlas readiness rubric.
IV-10 / WorkloadThe honest twin serverRead case file ↗Sources and provenance
- S-0015 / Tier B
Verifying LLM Inference to Detect Model Weight Exfiltration ↗
R. Rinberg, A. Karvonen, A. Hoover, D. Reuter, K. Warr · 2025 · arXiv
Supports: reference architecture; trust assumptions, including external control of user queries; benign evaluation prompts; adversarial prompts an open question; prefill-only cost; exfiltration results; stand-alone limitation
Locator: abstract; §4.2; §5; §5.3; §6.1; §8; Fig. 7
Version and catalogue details - S-0016 / Tier B
DiFR: Inference Verification Despite Nondeterminism ↗
A. Karvonen, D. Reuter, R. Rinberg, L. Marks, A. Garriga-Alonso, K. Warr · 2025 · ICML 2026 Workshop on Technical AI Governance Research
Supports: Token-DiFR and Activation-DiFR; benign nondeterminism; detection results; weights and sampling requirements; limitations and speculative-decoding sketch
Locator: abstract; §3.3; §5; §5.1; §7.2; §7.4; Appendix F
Version and catalogue details - S-0017 / Tier C
Example Schemes for Verifying High-Stakes AI Agreements ↗
Amodo Design · 2026 · Amodo Design
Supports: TOPLOC and Token-DiFR as example schemes; recomputation server in prover's data centre; correctness vs completeness
Locator: introduction; inference schemes
Version and catalogue details - S-0018 / Tier B
A System Overview for Near-Term, Low-Trust AI Compute Verification ↗
N. Cankaya · 2026 · Machine Intelligence Research Institute
Supports: bit-exact replay metadata; evaluation in auditing environment; sampling statistics; attribution problem
Locator: §2b; §3.2.2; §5.2.2; Appendix A1
Version and catalogue details - S-0020 / Tier B
Bit-Exact AI Inference Verification Without Performance Tradeoffs ↗
N. Cankaya · 2026 · ICML 2026 Workshop on Technical AI Governance Research
Supports: statistical schemes bound but do not close covert bandwidth; detection probability; bit-exact pass/fail
Locator: abstract; §1
Version and catalogue details - S-0067 / Tier C
Verification Plan ↗
R. Dean · 2026 · AI 2040
Supports: network taps feeding a recomputation server; partial recomputation of random samples
Locator: Concrete inference-only retrofitting proposal
Version and catalogue details - S-1000 / Tier A
TOPLOC: A Locality Sensitive Hashing Scheme for Trustless Verifiable Inference ↗
J. M. Ong, M. Di Ferrante, A. Pazdera, R. Garner, S. Jaghouar, M. Basra, M. Ryabinin, J. Hagemann · 2025 · Proceedings of the 42nd International Conference on Machine Learning (PMLR 267), pp. 47196-47211
Supports: TOPLOC mechanism, results and stated limitations, including untested KV-cache compression
Locator: abstract; §4; §5; §6.1-6.4
Version and catalogue details - S-1001 / Tier B
PrimeIntellect-ai/toploc (GitHub repository) ↗
Prime Intellect · 2025 · GitHub
Supports: public TOPLOC implementation
Locator: README; release v0.1.6
Version and catalogue details - S-1003 / Tier B
INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning ↗
Prime Intellect Team, S. Jaghouar, J. Mattern, J. M. Ong, J. Straube, M. Basra, A. Pazdera, K. Thaman, M. Di Ferrante, F. Gabriel, F. Obeid, K. Erdem, M. Keiblinger, J. Hagemann · 2025 · arXiv
Supports: developer use of TOPLOC to validate untrusted workers
Locator: §2.3; §2.4.2
Version and catalogue details - S-1005 / Tier B
adamkarvonen/difr (GitHub repository) ↗
A. Karvonen · 2025 · GitHub
Supports: public DiFR implementation with vLLM integration
Locator: README
Version and catalogue details - S-1006 / Tier C
Scaling Recomputation Inference Verification ↗
Amodo Design · 2026 · Amodo Design
Supports: prototype architecture, scale and verifier advantage
Locator: whole note; Fig. 3
Version and catalogue details - S-1007 / Tier B
Amodo-Design/Inference-Recomputation-Prototype (GitHub repository) ↗
Amodo Design · 2026 · GitHub
Supports: prototype code and workflow; DiFR module is a modified copy of the difr library
Locator: README
Version and catalogue details - S-3000 / Tier C
SYNTHETIC-2 Release: Four Million Collaboratively Generated Reasoning Traces ↗
Prime Intellect · 2025 · Prime Intellect blog
Supports: provider-reported TOPLOC v2 use in SYNTHETIC-2; 1,253 GPUs; false-positive rate over 4 million samples
Locator: verification section; GPU section
Version and catalogue details - S-3001 / Tier C
An Inference Verification Prototype — Stage 1 ↗
Amodo Design · 2026 · Amodo Design
Supports: Amodo's first prototype reuses the DiFR library; tests limited to a fraction of possible attacks
Locator: setup; limitations
Version and catalogue details - S-1008 / Tier C
AI 2040 Plan A — Verification SITREP ↗
Amodo Design · 2026 · Amodo Design
Supports: status of recomputation algorithms, capture, server security and red-teaming
Locator: status items
Version and catalogue details - S-1507 / Tier B
Adversarial Entropy Inflation Against Gumbel-Based Inference Verification ↗
N. Kezins · 2026 · arXiv
Supports: independent prompt-control attack on the weight-exfiltration detector; scope limited to the exfiltration bound
Locator: abstract; results
Version and catalogue details
- Source review date
- 2026-09-25
- Drafted by (source map)
- ai
- Review handles (source map)
- codex-review