M-0004 / Cryptography & computation

Zero-knowledge proofs of inference

A prover produces a cryptographic proof that an output came from running a committed model on a given input, without revealing the weights.

R2 DemonstratedSource reviewed 2026-09-25

01 / The mechanism and its boundary

What the technique establishes

Zero-knowledge proofs of inference let an AI developer show that a committed model computed an output from a given input, without disclosing the weights. A verifier checks a small proof in seconds. The peer-reviewed system zkLLM proved one 2,048-token forward pass of a 13-billion-parameter model in about 13 minutes on one A100 GPU. Its zero-knowledge guarantee assumes a verifier that follows the protocol, and its public research code is unaudited. A company, Attestable, reports proving a 31-billion-parameter model at 53 tokens per second for one 16,000-token sequence on one H100, without a paper or code. The main obstacle is cost. The main weaknesses are coverage and fidelity: a proof covers only the computation proven, and that computation is a fixed-point approximation of the model. Proofs also do not show how much computation produced an output. An independent audit of one proving library, ezkl, found high-severity soundness bugs, since fixed.

Threat model
Adversarial prover
Adversarial evaluation
Independent red-team
Hardware needed
None
Prover cooperation
Required
Confidentiality
Partial
Category
Cryptography & computation
Technical detail and cited results

Systems encode a network as an arithmetic circuit, or as sumcheck and lookup relations, over a finite field, with tensors scaled to fixed point S-0021 S-0023.

  • ZKML (EuroSys 2024) compiles models to halo2 circuits with KZG commitments (trusted setup) or IPA commitments (transparent), and optimises circuit layout. A distilled GPT-2 (81.3M parameters) took 3,651.67 s to prove with KZG, 18.70 s to verify and a 28,128-byte proof on a 128-vCPU, 1 TB machine S-0021.
  • ezkl also targets halo2 with KZG. South et al. proved a 250,000-parameter nanoGPT in 2,781 s, verified it in 2.69 s and needed a 219 GB proving key S-0024. A Trail of Bits audit of ezkl (January 2025, commit bdcba5c) reported 34 findings, 8 of high severity; three high-severity findings were circuit soundness issues, and all high-severity findings were resolved at the March 2025 fix review S-0070.
  • zkLLM (CCS 2024) uses sumcheck-based arguments, a parallel lookup argument (tlookup) for non-arithmetic tensor operations, an attention protocol (zkAttn) and Hyrax commitments over BLS12-381, with values scaled by 2^16. On one A100 40 GB GPU with 2,048-token inputs, LLaMa-2-13B needed 986 s for the one-time weight commitment, 803 s to prove, 188 kB of proof, 3.95 s to verify and 23.1 GB of memory. Perplexity on C4 moved from 6.520 to 6.528 S-0023.
  • zkGPT (USENIX Security 2025) combines the GKR protocol, Lasso lookups and Hyrax commitments, and is made non-interactive with Fiat–Shamir. On a 16-core Xeon server with 200 GB of memory, it proved GPT-2 inference in 21.8 s with 32 threads, with a 101 KB proof verified in 0.35 s. Its code is archived on Zenodo S-3060.
  • NanoZK (ICICS 2026) proves each transformer layer separately with Halo2/IPA over the Pallas curve, uses 16-bit lookup tables for softmax, GELU and normalisation, and chains layer boundaries with SHA-256 commitments. It reports 3.2 to 3.7 KB per sub-circuit (attention or MLP) proof, about 83 KB in total for 12 layers. It measured full-block proofs on CPU up to width 128 and, assuming a GPU speedup, projects about 68 s per block at GPT-2 width, or about 14 minutes to prove a 12-layer GPT-2 sequentially S-0068.
  • Attestable reports a prover whose security rests only on hash functions, at 100-bit security, with 8-bit integer matrix multiplications. On one H100 it reports 4.35 to 7.92 MiB proofs, 157 to 648 ms CPU verification, and 53 tokens per second for one 16K-token sequence of a 31B Gemma model S-1101.

Claims and scope

A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.

Readiness for a stated use

R2 Demonstrated

Assessed use: proving a language model's output follows from committed weights, against a cheating prover

medium confidence · current · assessed 2026-10-08 · rubric 1.1

This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.

R2 through zkLLM, which has public, artifact-evaluated code and peer-reviewed results at 13 billion parameters; no implementation is yet production-grade for language models.

  • R1 met: ZKML, zkLLM and NanoZK describe the claim (an output equals the committed model applied to an input) and their assumptions S-0021 S-0023 S-0068.
  • R2 met through zkLLM. Its code is public and received CCS 2024 artifact-evaluation badges S-1108. Its published end-to-end results use a 13-billion-parameter model on a data-centre GPU, against a stated adversary: a cheating polynomial-time prover S-0023.
  • R3 not met for this use. zkLLM's README says the code is not ready for industrial applications S-1108. Attestable's results come without public code, paper or reproducible artifacts S-1101. The ezkl library is public, and Trail of Bits reports that other projects use its verifier contracts in production S-0070. The language model South et al. prove with ezkl is a 250,000-parameter nanoGPT, and the largest model in their results is a VAE decoder of about 1.07 million parameters S-0024. Both are far below the scale this use concerns. Lagrange calls DeepProve "production-ready" S-1807, but its repository has no releases, and its reported benchmarks are GPT-2 and Gemma 3 at 512 tokens S-1808. No source documents a party other than a developer relying on any of these systems for a verification decision about a language model.
  • R4 not met. Only ezkl has an independent audit: it left no high-severity finding unresolved S-0070. As of September 2026 no independent audit or red-team of zkLLM has been published. One independent analysis, tested on a small transformer, shows that valid proofs of LLM inference do not bind the computation spent, so a much smaller model can pass as the declared one S-1112.

Evidence needed for the next level

  • An implementation that is production-grade and available for language models, or relied on by a party other than its developer for a verification decision.

  • An independent public security evaluation (audit, red-team or peer-reviewed analysis) of a proof system that handles language models.

Limitations, flaws, and blockers

These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.

significant / open / open question

The proof covers a fixed-point approximation, not the floating-point model

Current ZK inference systems prove a quantised version of the network. zkLLM scales values by 2^16 and reports small perplexity changes S-0023. Attestable reports quantising matrix multiplications to 8-bit integers while proving other operations in floating point S-1101. A verifier therefore learns about the proof-friendly variant, and must separately accept that this variant is the declared model. Trail of Bits built a ResNet-18 backdoor that is dormant in the full-precision model and active after ezkl's quantisation; whether it persists through proving was left for further investigation S-0070. A verification system design calls floating-point emulation in ZKPs an open problem S-0018.

S-0023S-1101S-0070S-0018

significant / open / theoretical argument

A proof speaks only for the computations that were proven

Attestable writes that "a proof of some computation is not a proof of all computation", and that a proof cannot discover a datacenter that was never declared S-1102. Proofs of inference do not by themselves show that no other workload ran on the same or other hardware.

S-1102

minor / open / theoretical argument

The model architecture is disclosed

ZKML "requires that the model architecture (but not weights) is revealed" S-0021, and zkLLM assumes a publicly known model structure S-0023. Architecture can be commercially sensitive.

S-0021S-0023

significant / open / demonstrated attack

Proofs do not bind computational effort (Hollow-LLM)

Researchers at the University of Southern California show that a proof of inference certifies that an output is consistent with committed weights under the declared architecture, but not how much computation produced it S-1112. In their Hollow-LLM attack, a provider keeps the declared architecture and parameter count but commits to "ghost weights". Some layers pass their inputs through unchanged, and wide layers carry the signal in a small subspace, so a much smaller inner model does the real work. The ghost weights satisfy the verification circuit and yield valid proofs S-1112.

The authors ran the attack with the proof procedure of zkGPT, a separate ZK inference system, on a 6-layer, 512-dimensional transformer declared as up to 12 layers and 1,024 dimensions. Outputs were identical to the inner model's, and serving cost stayed at the inner model's level. An honest model of the declared size cost 2.4 times as much to prefill and 3.1 times as much to decode. Proving cost still grew with the declared architecture S-1112.

The authors note that results may be served before any proof, with the provider building the witness only when a call is selected for audit. They describe their constructions as "compatible with state-of-the-art zkLLM pipelines", and state that the attack does not imply a flaw in the proof system itself. They propose challenge-based audits and ablation tests, which raise the cost of cheating but give no guarantee S-1112.

S-1112

What still blocks use or stronger assurance

  1. Proving takes about 13 minutes (803 seconds) per 2,048-token forward pass of a 13B model on one A100 S-0023, and a verification system design calls the overhead heavy S-0018.

    S-0023S-0018
  2. ZKML and zkLLM prove fixed-point arithmetic S-0021 S-0023, and floating-point emulation in ZKPs is described as an open problem S-0018.

    S-0021S-0023S-0018
  3. zkLLM's code is unaudited, interactive and archived S-1108; the one audited ZK inference library, ezkl, had high-severity circuit soundness bugs before its fixes S-0070.

    S-1108S-0070
  4. Showing that proven inference was the only work done needs a compute-accounting mechanism such as proof-of-work accounting, which is only proposed S-1102.

    Dependency: Proofs of useful work for capacity accounting

    S-1102

Connections in the research map

Complementary techniques

Alternative approaches

Concepts used

Organizations and developers

Implementations

Sources and provenance

  1. S-0023 / Tier A

    zkLLM: Zero Knowledge Proofs for Large Language Models ↗

    H. Sun, J. Li, H. Zhang · 2024 · 2024 ACM SIGSAC Conference on Computer and Communications Security (CCS 2024)

    Supports: zkLLM design, threat model, security theorems, overheads, fixed-point effects

    Locator: abstract; §3.6; §4–5; §7.2 Theorems 7.3–7.4; §8 Table 1; §9

    Version and catalogue details
  2. S-1108 / Tier B

    zkllm-ccs2024: code for zkLLM: Zero Knowledge Proofs for Large Language Models ↗

    H. Sun · 2024 · GitHub; archived on Zenodo

    Supports: zkLLM code availability, artifact badges and README caveats

    Locator: README; Zenodo record

    Version and catalogue details
  3. S-0021 / Tier A

    ZKML: An Optimizing System for ML Inference in Zero-Knowledge Proofs ↗

    B.-J. Chen, S. Waiwitlikhit, I. Stoica, D. Kang · 2024 · 19th European Conference on Computer Systems (EuroSys 2024)

    Supports: ZKML design, halo2 backends, GPT-2 overheads, limitations

    Locator: §3; §4.1; §4.4; §9 Tables 5–7

    Version and catalogue details
  4. S-0068 / Tier A

    NanoZK: Privacy-Preserving Verifiable Inference for Large Language Models via Layerwise Zero-Knowledge Proofs ↗

    Z. Wang · 2026 · International Conference on Information and Communications Security (ICICS 2026)

    Supports: NanoZK layerwise proofs, threat model, proof sizes, partial audits

    Locator: Definition 1; §5; §6; Table 8; App. A.4

    Version and catalogue details
  5. S-0024 / Tier B

    Verifiable evaluations of machine learning models using zkSNARKs ↗

    T. South, A. Camuto, S. Jain, S. Nguyen, R. Mahari, C. Paquin, J. Morton, A. Pentland · 2024 · arXiv

    Supports: verifiable evaluation attestations with ezkl; public inputs and outputs; costs of small models

    Locator: abstract; §5; §6.1 Table 1

    Version and catalogue details
  6. S-0070 / Tier B

    Zkonduit EZKL Security Assessment ↗

    F. Casal, T. Hess, L. Bourtoule, S. Hussain, G. Larregay · 2025 · Trail of Bits (prepared for Zkonduit Inc.)

    Supports: independent audit of ezkl: circuit soundness findings, quantisation-activated backdoor, production use, fix review

    Locator: Executive Summary; findings TOB-EZKL-4 to 6 and 17; App. D

    Version and catalogue details
  7. S-1100 / Tier A

    A Survey of Zero-Knowledge Proof Based Verifiable Machine Learning ↗

    Z. Peng, C. Zhao, T. Wang, G. Liao, Z. Lin, Y. Liu, B. Cao, L. Shi, Q. Yang, S. Zhang · 2026 · Artificial Intelligence Review, vol. 59, no. 7, article 157

    Supports: definition and categorisation of ZKML; main implementation bottlenecks

    Locator: abstract; §III; Table VI

    Version and catalogue details
  8. S-1101 / Tier C

    Proving LLMs at Scale ↗

    Attestable · 2026 · Attestable blog

    Supports: Attestable's reported prover, statement proven, performance and limits (provider-reported)

    Version and catalogue details
  9. S-1102 / Tier C

    Pacing AI Requires Proof ↗

    Attestable · 2026 · Attestable blog

    Supports: Attestable's coverage argument and pacing proposal (provider-reported)

    Version and catalogue details
  10. S-1103 / Tier C

    From Verifiability to Model-Weight Security ↗

    Attestable · 2026 · Attestable blog

    Supports: Attestable's proposal to prove randomly sampled outputs (provider-reported)

    Version and catalogue details
  11. S-0018 / Tier B

    A System Overview for Near-Term, Low-Trust AI Compute Verification ↗

    N. Cankaya · 2026 · Machine Intelligence Research Institute

    Supports: ZKPs as a 'tentative plan B' in a verification system; overhead assessment; floating-point gap

    Locator: §5.2.4

    Version and catalogue details
  12. S-1112 / Tier B

    Hollow-LLM Attack: Computationally Trivial Weights in Zero-Knowledge Verification of LLM Inference ↗

    C. Gong, B. Liu, M. Li · 2026 · arXiv

    Supports: independent Hollow-LLM analysis: proofs do not bind computational effort; ghost-weight constructions; zkGPT-based experiment and cost results; countermeasures

    Locator: Abstract; §I contributions; §V Table 2; §VI

    Version and catalogue details
  13. S-3060 / Tier A

    zkGPT: An Efficient Non-interactive Zero-knowledge Proof Framework for LLM Inference ↗

    W. Qu, Y. Sun, X. Liu, T. Lu, Y. Guo, K. Chen, J. Zhang · 2025 · 34th USENIX Security Symposium (USENIX Security 25), pp. 2045–2063

    Supports: zkGPT design, non-interactive proofs, GPT-2 proving and verification figures, public code

    Locator: abstract; §3; §6 Table 3

    Version and catalogue details
  14. S-1807 / Tier C

    DeepProve-1: The First zkML System to Prove a Full LLM Inference ↗

    Lagrange Labs · 2025 · Lagrange blog

    Supports: Lagrange's reported proof of full GPT-2 inference and 'production-ready' description (provider-reported)

    Version and catalogue details
  15. S-1808 / Tier B

    Lagrange-Labs/deep-prove (GitHub repository) ↗

    Lagrange Labs · 2026 · GitHub

    Supports: DeepProve public code, licence and reported GPT-2 proving figures (provider-reported)

    Locator: README

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review