I-0001 / Cryptography & computation

TOPLOC

TOPLOC is a hashing scheme from Prime Intellect that lets a verifier check whether an inference provider ran the model, prompt and precision it claims.

R3 In productionSource reviewed 2026-09-25Provider-reported evidence

01 / The mechanism and its boundary

What the technique establishes

TOPLOC is a hashing scheme for checking that an inference provider ran the model, prompt and numerical precision it claims. During generation, the provider records a compact fingerprint of the model's largest last-layer activations. A verifier re-runs the sequence in one pass and checks that the fingerprints match within set tolerances. The tolerances absorb harmless differences between GPUs. The peer-reviewed paper reports catching every tested change to model, prompt or precision, with no false positives or negatives. Proofs take 258 bytes per 32 generated tokens. TOPLOC is open source. Prime Intellect, its developer, used it in 2025 to accept or reject work from untrusted computers in decentralized training and data-generation runs. As of September 2026 no independent security evaluation has been published. The authors list attacks it cannot yet catch, such as speculative decoding with a cheaper model. Subtle changes are also harder to detect than large ones.

Threat model
Adversarial prover
Adversarial evaluation
Published analysis
Hardware needed
None
Prover cooperation
Required
Confidentiality
Revealing
Category
Cryptography & computation
Technical detail and cited results
  • The prover commits to its activations every 32 generated tokens. It takes the top-k values of the last hidden layer, with k = 128 in the main configuration. It encodes their indices and values as a polynomial over an integer field, with a modulus chosen to be injective on the index set S-1000. The result is k two-byte coefficients. For Llama 3.1-8B-Instruct that is 258 bytes per 32 tokens, against 262 KB for storing the embeddings directly S-1000.
  • The verifier decodes the proof and recomputes the top-k values with a prefill pass. It counts exponent mismatches and computes the mean and median mantissa differences. Validation succeeds if all three are below their thresholds. For bf16 the thresholds are 38, 10 and 8 S-1000.
  • The hardware tests used 1× A100, 1× RTX 4090 and 2× RTX 4090 GPUs, with FlashAttention 2, PyTorch SDPA and FlexAttention. The authors read the activations through a vLLM hook S-1000.
  • Prime Intellect reports that validation is up to 100 times faster than the original inference S-1002 S-1003. In SYNTHETIC-2 it reports a median verification cost, averaged across models, 25 times lower than re-running the inference S-3000. It reports that proof generation cut tokens-per-second throughput by about 1% in INTELLECT-2 S-1003.
  • Prime Intellect reports that TOPLOC v2 adds reproducible Gumbel noise for categorical sampling, so that verifiers can check token sampling. Version 2 also extends the scheme to pipeline-parallel inference. A pass at the final pipeline stage accepts all stages, and a failure triggers a stage-by-stage replay that finds the first faulty node S-1004 S-3000.
  • The package is published on PyPI as toploc. The latest tag is v0.1.6 S-1001.

Claims and scope

A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.

Readiness for a stated use

R3 In production

Assessed use: checking that untrusted providers used the claimed model, prompt and precision

low confidence · current · assessed 2026-09-25 · rubric 1.1

This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.

Prime Intellect, its developer, has used the public package in production to accept or reject work from untrusted computers, but no independent security audit or red-team has been published.

  • R1 met: the peer-reviewed paper sets out the design, the claim and the threat S-1000. The claim is that the provider used the stated model, prompt and precision. The threat is undisclosed changes to any of them.
  • R2 met: a public MIT-licensed implementation exists S-1001. The paper reports results on A100 and RTX 4090 GPUs across several models, attention implementations and one- and two-GPU tensor parallelism S-1000.
  • R3 met through the developer's own use: the package is public S-1001, and Prime Intellect reports using it in production to accept or reject work from untrusted inference workers, in a 32-billion-parameter decentralized training run and a data-generation run on 1,253 GPUs S-1003 S-3000. The latest documented use is the data-generation run, whose results Prime Intellect released in July 2025 S-3000. No other party is documented relying on TOPLOC for a verification decision.
  • R4 not met: no independent audit, red-team or peer-reviewed security analysis has been published. DiFR's comparison measures detection accuracy against communication cost S-0016. It is not a security evaluation.

Confidence is low: the production use is the developer's own, and none is documented after July 2025. The package has had no release since April 2025 S-1001.

Evidence needed for the next level

  • An independent public security evaluation, such as an audit, red-team or peer-reviewed analysis, that tests adaptive attacks like the spoofing and speculative-decoding cases the TOPLOC authors list.

Limitations, flaws, and blockers

These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.

significant / open / theoretical argument

Speculative decoding goes undetected

The TOPLOC authors state that it cannot detect speculative decoding. In speculative decoding, a provider decodes with a cheaper model and uses the larger model only for prefill.

S-1000

significant / open / open question

Last-layer activations could be spoofed

The TOPLOC authors name spoofing of the last hidden layer's activations as a potential attack. A provider could do this by pruning intermediate layers or by using a smaller model.

S-1000

significant / open / open question

Subtle modifications are harder to detect

The TOPLOC authors state that large changes to the model or prompt are straightforward to detect, but subtle modifications are harder. In preliminary experiments, the margin separating fp8 from bf16 generation was small. The authors did not test whether TOPLOC distinguishes types of KV-cache compression.

S-1000

significant / open / theoretical argument

Tolerance leaves covert bandwidth

TOPLOC accepts approximate matches. A check of this kind can put an upper bound on the covert bandwidth available to an adversary, but it cannot close that bandwidth. The limit applies to all statistical verification schemes.

S-0020

What still blocks use or stronger assurance

  1. No independent security evaluation has been published, and Amodo Design rates red-teaming of recomputation schemes as 'not started'.

    S-1008
  2. The verifier must run the model itself, which suits the paper's setting of providers serving open-weights models.

    S-1000

Connections in the research map

Mechanisms implemented

Concepts used

Organizations and developers

Sources and provenance

  1. S-1000 / Tier A

    TOPLOC: A Locality Sensitive Hashing Scheme for Trustless Verifiable Inference ↗

    J. M. Ong, M. Di Ferrante, A. Pazdera, R. Garner, S. Jaghouar, M. Basra, M. Ryabinin, J. Hagemann · 2025 · Proceedings of the 42nd International Conference on Machine Learning (PMLR 267), pp. 47196-47211

    Supports: design, commitment and validation algorithm, thresholds, experiments, limitations

    Locator: abstract; §3.1; §4; §5.1-5.7; §6.1-6.5

    Version and catalogue details
  2. S-1001 / Tier B

    PrimeIntellect-ai/toploc (GitHub repository) ↗

    Prime Intellect · 2025 · GitHub

    Supports: public implementation, licence, release tag

    Locator: README; releases

    Version and catalogue details
  3. S-1002 / Tier C

    TOPLOC: A Locality Sensitive Hashing Scheme for Trustless Verifiable Inference (blog post) ↗

    Prime Intellect · 2025 · Prime Intellect blog

    Supports: provider-reported validation speed and SGLang/vLLM integrations

    Locator: whole post

    Version and catalogue details
  4. S-1003 / Tier B

    INTELLECT-2: A Reasoning Model Trained Through Globally Decentralized Reinforcement Learning ↗

    Prime Intellect Team, S. Jaghouar, J. Mattern, J. M. Ong, J. Straube, M. Basra, A. Pazdera, K. Thaman, M. Di Ferrante, F. Gabriel, F. Obeid, K. Erdem, M. Keiblinger, J. Hagemann · 2025 · arXiv

    Supports: provider-reported use in INTELLECT-2; checks; eviction of failing nodes; proof-generation overhead; validation speed

    Locator: §2.3; §2.4.2

    Version and catalogue details
  5. S-1004 / Tier C

    SYNTHETIC-2 ↗

    Prime Intellect · 2025 · Prime Intellect blog

    Supports: provider-reported TOPLOC v2 sampling verification, pipeline-parallel extension and pipeline replay in SYNTHETIC-2

    Locator: verification section

    Version and catalogue details
  6. S-3000 / Tier C

    SYNTHETIC-2 Release: Four Million Collaboratively Generated Reasoning Traces ↗

    Prime Intellect · 2025 · Prime Intellect blog

    Supports: provider-reported TOPLOC v2 use in SYNTHETIC-2: group-level accept/reject, stage-by-stage replay, sampling proofs, false-positive rate, verification cost, 1,253 GPUs

    Locator: verification section; GPU section

    Version and catalogue details
  7. S-0016 / Tier B

    DiFR: Inference Verification Despite Nondeterminism ↗

    A. Karvonen, D. Reuter, R. Rinberg, L. Marks, A. Garriga-Alonso, K. Warr · 2025 · ICML 2026 Workshop on Technical AI Governance Research

    Supports: independent comparison with Activation-DiFR on detecting FP8 KV-cache quantization

    Locator: §6.2; Table 1

    Version and catalogue details
  8. S-0017 / Tier C

    Example Schemes for Verifying High-Stakes AI Agreements ↗

    Amodo Design · 2026 · Amodo Design

    Supports: independent description of the TOPLOC scheme

    Locator: TOPLOC section

    Version and catalogue details
  9. S-0020 / Tier B

    Bit-Exact AI Inference Verification Without Performance Tradeoffs ↗

    N. Cankaya · 2026 · ICML 2026 Workshop on Technical AI Governance Research

    Supports: limits of statistical verification

    Locator: §1

    Version and catalogue details
  10. S-1008 / Tier C

    AI 2040 Plan A — Verification SITREP ↗

    Amodo Design · 2026 · Amodo Design

    Supports: status as an initial recomputation scheme under testing; recomputation red-teaming rated 'not started'

    Locator: Recomputation algorithms and Recomputation red-teaming items

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review