A proposed retrofit that isolates data-centre inference units, taps their front-end traffic and recomputes random samples to check that only declared inference runs.
Assessed use: showing that retrofitted data centres run only inference
Apple's cloud AI inference service, in which user devices send requests only to servers that attest to running software published in a public transparency log.
Assessed use: showing users which software serves their AI requests, not which model
A research prototype that runs AI safety benchmarks inside a trusted execution environment and publishes attestations binding the model, the audit and the results.
Assessed use: showing users that the model answering them is the audited one
Attestable's zero-knowledge prover, which the company reports proves large language model outputs came from committed weights at tens of tokens per second.
Assessed use: proving an output came from committed weights
Capping or removing network links between groups of accelerators, so that serving within each group still works but large training across groups becomes far slower.
Assessed use: monitoring inter-node traffic with operator-run software on four GPUs
Open-source kernels from Thinking Machines Lab that make LLM outputs independent of batch size, adopted in vLLM and SGLang to give reproducible inference.
Assessed use: exact recomputation of served outputs by a verifier, with a cooperating provider
Recording each AI chip's identity and owner from the fab onwards, and cryptographically fixing manufacturing records, so that chips can be accounted for later.
Assessed use: a checkable record of which chips were made and who declared owning them
Lets mutually distrusting parties run an agreed check over private models or records inside attested enclaves or zero-knowledge proofs, revealing only the result.
Assessed use: audits or evaluations of a private model that reveal neither party's inputs
DiFR checks that an inference provider ran its declared model by comparing output tokens or activations with a trusted re-run using the same random seed.
Assessed use: checking that outputs match the declared model, precision and sampling settings
Proposed chip add-ons, a guarantee processor inside a tamper-protected enclosure, that would check and enforce agreed rules on how AI accelerators are used.
Assessed use: checking and enforcing training-compute limits on chips, against adversaries up to states
A retrofittable reference design in which network taps commit to all facility traffic, and air-gapped, independently sourced checkers later re-run randomly challenged records.
Assessed use: screening challenged records to show declared inference compute is not training
A draft specification, hosted by Lucid Computing, for short-lived certificates that bound where a workload runs by timing signed exchanges with fixed anchors.
Assessed use: certifying the region where an attested workload ran at a given time
Cryptographic evidence that a given amount of matrix-multiplication work was completed, proposed as one input to accounting for spare capacity on declared hardware.
Assessed use: bounding the spare capacity of declared hardware that could run training
A RAND design for a purpose-built facility that serves already-trained AI models while protecting weights and inference data against state-level attackers.
Assessed use: the operator's own weight security, with no outside verification described
Remote detection locates large data centres and estimates their power capacity without site access, using satellite imagery, heat signatures and public records such as permits.
Assessed use: finding undeclared data centres above an agreed compute threshold
An open-source prototype that routes a facility's inference traffic through a logger and re-runs requests on a separate cluster to check it serves inference.
Assessed use: telling inference from training on a mutually inspected cluster
Tinfoil's method for proving which model weights its enclave-hosted inference service runs, by binding a dm-verity hash of the weights into remote attestation.
Assessed use: showing clients that the served weights match a committed hash
TOPLOC is a hashing scheme from Prime Intellect that lets a verifier check whether an inference provider ran the model, prompt and precision it claims.
Assessed use: checking that untrusted providers used the claimed model, prompt and precision
Gensyn's system for checking delegated machine-learning jobs, which settles disagreements between providers by re-running a single operation with bitwise-reproducible operators.
Assessed use: reproducing declared-model inference from receipts in Gensyn's information-market service
zkLLM is a GPU-accelerated zero-knowledge proof system that proves a large language model's output came from committed weights without revealing those weights.
Assessed use: proving an output came from committed weights, against a prover who cheats