01 / The mechanism and its boundary
What the technique establishes
Zero-knowledge proofs of inference let an AI developer show that a committed model computed an output from a given input, without disclosing the weights. A verifier checks a small proof in seconds. The peer-reviewed system zkLLM proved one 2,048-token forward pass of a 13-billion-parameter model in about 13 minutes on one A100 GPU. Its zero-knowledge guarantee assumes a verifier that follows the protocol, and its public research code is unaudited. A company, Attestable, reports proving a 31-billion-parameter model at 53 tokens per second for one 16,000-token sequence on one H100, without a paper or code. The main obstacle is cost. The main weaknesses are coverage and fidelity: a proof covers only the computation proven, and that computation is a fixed-point approximation of the model. Proofs also do not show how much computation produced an output. An independent audit of one proving library, ezkl, found high-severity soundness bugs, since fixed.
- Threat model
- Adversarial prover
- Adversarial evaluation
- Independent red-team
- Hardware needed
- None
- Prover cooperation
- Required
- Confidentiality
- Partial
- Category
- Cryptography & computation
Technical detail and cited results
Systems encode a network as an arithmetic circuit, or as sumcheck and lookup relations, over a finite field, with tensors scaled to fixed point S-0021 S-0023.
- ZKML (EuroSys 2024) compiles models to halo2 circuits with KZG commitments (trusted setup) or IPA commitments (transparent), and optimises circuit layout. A distilled GPT-2 (81.3M parameters) took 3,651.67 s to prove with KZG, 18.70 s to verify and a 28,128-byte proof on a 128-vCPU, 1 TB machine S-0021.
- ezkl also targets halo2 with KZG. South et al. proved a 250,000-parameter nanoGPT in 2,781 s, verified it in 2.69 s and needed a 219 GB proving key S-0024. A Trail of Bits audit of ezkl (January 2025, commit bdcba5c) reported 34 findings, 8 of high severity; three high-severity findings were circuit soundness issues, and all high-severity findings were resolved at the March 2025 fix review S-0070.
- zkLLM (CCS 2024) uses sumcheck-based arguments, a parallel lookup argument (tlookup) for non-arithmetic tensor operations, an attention protocol (zkAttn) and Hyrax commitments over BLS12-381, with values scaled by 2^16. On one A100 40 GB GPU with 2,048-token inputs, LLaMa-2-13B needed 986 s for the one-time weight commitment, 803 s to prove, 188 kB of proof, 3.95 s to verify and 23.1 GB of memory. Perplexity on C4 moved from 6.520 to 6.528 S-0023.
- zkGPT (USENIX Security 2025) combines the GKR protocol, Lasso lookups and Hyrax commitments, and is made non-interactive with Fiat–Shamir. On a 16-core Xeon server with 200 GB of memory, it proved GPT-2 inference in 21.8 s with 32 threads, with a 101 KB proof verified in 0.35 s. Its code is archived on Zenodo S-3060.
- NanoZK (ICICS 2026) proves each transformer layer separately with Halo2/IPA over the Pallas curve, uses 16-bit lookup tables for softmax, GELU and normalisation, and chains layer boundaries with SHA-256 commitments. It reports 3.2 to 3.7 KB per sub-circuit (attention or MLP) proof, about 83 KB in total for 12 layers. It measured full-block proofs on CPU up to width 128 and, assuming a GPU speedup, projects about 68 s per block at GPT-2 width, or about 14 minutes to prove a 12-layer GPT-2 sequentially S-0068.
- Attestable reports a prover whose security rests only on hash functions, at 100-bit security, with 8-bit integer matrix multiplications. On one H100 it reports 4.35 to 7.92 MiB proofs, 157 to 648 ms CPU verification, and 53 tokens per second for one 16K-token sequence of a 31B Gemma model S-1101.
Claims and scope
A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.
The declared model is the one being served
Binds each proven output to committed weights and a public architecture.
This compute runs inference, not training
Attestable proposes using proofs to show accounted workloads used an approved, unchanged model.
Declared safeguards were applied during inference
Attestable proposes that a proof could show an agreed input classifier was applied; South et al. prove evaluation results.
Readiness for a stated use
Assessed use: proving a language model's output follows from committed weights, against a cheating prover
medium confidence · current · assessed 2026-10-08 · rubric 1.1
This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.
R2 through zkLLM, which has public, artifact-evaluated code and peer-reviewed results at 13 billion parameters; no implementation is yet production-grade for language models.
- R1 met: ZKML, zkLLM and NanoZK describe the claim (an output equals the committed model applied to an input) and their assumptions S-0021 S-0023 S-0068.
- R2 met through zkLLM. Its code is public and received CCS 2024 artifact-evaluation badges S-1108. Its published end-to-end results use a 13-billion-parameter model on a data-centre GPU, against a stated adversary: a cheating polynomial-time prover S-0023.
- R3 not met for this use. zkLLM's README says the code is not ready for industrial applications S-1108. Attestable's results come without public code, paper or reproducible artifacts S-1101. The ezkl library is public, and Trail of Bits reports that other projects use its verifier contracts in production S-0070. The language model South et al. prove with ezkl is a 250,000-parameter nanoGPT, and the largest model in their results is a VAE decoder of about 1.07 million parameters S-0024. Both are far below the scale this use concerns. Lagrange calls DeepProve "production-ready" S-1807, but its repository has no releases, and its reported benchmarks are GPT-2 and Gemma 3 at 512 tokens S-1808. No source documents a party other than a developer relying on any of these systems for a verification decision about a language model.
- R4 not met. Only ezkl has an independent audit: it left no high-severity finding unresolved S-0070. As of September 2026 no independent audit or red-team of zkLLM has been published. One independent analysis, tested on a small transformer, shows that valid proofs of LLM inference do not bind the computation spent, so a much smaller model can pass as the declared one S-1112.
Evidence needed for the next level
An implementation that is production-grade and available for language models, or relied on by a party other than its developer for a verification decision.
An independent public security evaluation (audit, red-team or peer-reviewed analysis) of a proof system that handles language models.
Limitations, flaws, and blockers
These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.
significant / open / open question
The proof covers a fixed-point approximation, not the floating-point model
Current ZK inference systems prove a quantised version of the network. zkLLM scales values by 2^16 and reports small perplexity changes S-0023. Attestable reports quantising matrix multiplications to 8-bit integers while proving other operations in floating point S-1101. A verifier therefore learns about the proof-friendly variant, and must separately accept that this variant is the declared model. Trail of Bits built a ResNet-18 backdoor that is dormant in the full-precision model and active after ezkl's quantisation; whether it persists through proving was left for further investigation S-0070. A verification system design calls floating-point emulation in ZKPs an open problem S-0018.
significant / open / theoretical argument
A proof speaks only for the computations that were proven
Attestable writes that "a proof of some computation is not a proof of all computation", and that a proof cannot discover a datacenter that was never declared S-1102. Proofs of inference do not by themselves show that no other workload ran on the same or other hardware.
minor / open / theoretical argument
The model architecture is disclosed
ZKML "requires that the model architecture (but not weights) is revealed" S-0021, and zkLLM assumes a publicly known model structure S-0023. Architecture can be commercially sensitive.
significant / open / demonstrated attack
Proofs do not bind computational effort (Hollow-LLM)
Researchers at the University of Southern California show that a proof of inference certifies that an output is consistent with committed weights under the declared architecture, but not how much computation produced it S-1112. In their Hollow-LLM attack, a provider keeps the declared architecture and parameter count but commits to "ghost weights". Some layers pass their inputs through unchanged, and wide layers carry the signal in a small subspace, so a much smaller inner model does the real work. The ghost weights satisfy the verification circuit and yield valid proofs S-1112.
The authors ran the attack with the proof procedure of zkGPT, a separate ZK inference system, on a 6-layer, 512-dimensional transformer declared as up to 12 layers and 1,024 dimensions. Outputs were identical to the inner model's, and serving cost stayed at the inner model's level. An honest model of the declared size cost 2.4 times as much to prefill and 3.1 times as much to decode. Proving cost still grew with the declared architecture S-1112.
The authors note that results may be served before any proof, with the provider building the witness only when a call is selected for audit. They describe their constructions as "compatible with state-of-the-art zkLLM pipelines", and state that the attack does not imply a flaw in the proof system itself. They propose challenge-based audits and ablation tests, which raise the cost of cheating but give no guarantee S-1112.
What still blocks use or stronger assurance
- S-0023S-0018
Proving takes about 13 minutes (803 seconds) per 2,048-token forward pass of a 13B model on one A100 S-0023, and a verification system design calls the overhead heavy S-0018.
- S-0021S-0023S-0018
ZKML and zkLLM prove fixed-point arithmetic S-0021 S-0023, and floating-point emulation in ZKPs is described as an open problem S-0018.
- S-1108S-0070
zkLLM's code is unaudited, interactive and archived S-1108; the one audited ZK inference library, ezkl, had high-severity circuit soundness bugs before its fixes S-0070.
Showing that proven inference was the only work done needs a compute-accounting mechanism such as proof-of-work accounting, which is only proposed S-1102.
Dependency: Proofs of useful work for capacity accounting
S-1102
Connections in the research map
Complementary techniques
Alternative approaches
Concepts used
Organizations and developers
Implementations
Sources and provenance
- S-0023 / Tier A
zkLLM: Zero Knowledge Proofs for Large Language Models ↗
H. Sun, J. Li, H. Zhang · 2024 · 2024 ACM SIGSAC Conference on Computer and Communications Security (CCS 2024)
Supports: zkLLM design, threat model, security theorems, overheads, fixed-point effects
Locator: abstract; §3.6; §4–5; §7.2 Theorems 7.3–7.4; §8 Table 1; §9
Version and catalogue details - S-1108 / Tier B
zkllm-ccs2024: code for zkLLM: Zero Knowledge Proofs for Large Language Models ↗
H. Sun · 2024 · GitHub; archived on Zenodo
Supports: zkLLM code availability, artifact badges and README caveats
Locator: README; Zenodo record
Version and catalogue details - S-0021 / Tier A
ZKML: An Optimizing System for ML Inference in Zero-Knowledge Proofs ↗
B.-J. Chen, S. Waiwitlikhit, I. Stoica, D. Kang · 2024 · 19th European Conference on Computer Systems (EuroSys 2024)
Supports: ZKML design, halo2 backends, GPT-2 overheads, limitations
Locator: §3; §4.1; §4.4; §9 Tables 5–7
Version and catalogue details - S-0068 / Tier A
NanoZK: Privacy-Preserving Verifiable Inference for Large Language Models via Layerwise Zero-Knowledge Proofs ↗
Z. Wang · 2026 · International Conference on Information and Communications Security (ICICS 2026)
Supports: NanoZK layerwise proofs, threat model, proof sizes, partial audits
Locator: Definition 1; §5; §6; Table 8; App. A.4
Version and catalogue details - S-0024 / Tier B
Verifiable evaluations of machine learning models using zkSNARKs ↗
T. South, A. Camuto, S. Jain, S. Nguyen, R. Mahari, C. Paquin, J. Morton, A. Pentland · 2024 · arXiv
Supports: verifiable evaluation attestations with ezkl; public inputs and outputs; costs of small models
Locator: abstract; §5; §6.1 Table 1
Version and catalogue details - S-0070 / Tier B
Zkonduit EZKL Security Assessment ↗
F. Casal, T. Hess, L. Bourtoule, S. Hussain, G. Larregay · 2025 · Trail of Bits (prepared for Zkonduit Inc.)
Supports: independent audit of ezkl: circuit soundness findings, quantisation-activated backdoor, production use, fix review
Locator: Executive Summary; findings TOB-EZKL-4 to 6 and 17; App. D
Version and catalogue details - S-1100 / Tier A
A Survey of Zero-Knowledge Proof Based Verifiable Machine Learning ↗
Z. Peng, C. Zhao, T. Wang, G. Liao, Z. Lin, Y. Liu, B. Cao, L. Shi, Q. Yang, S. Zhang · 2026 · Artificial Intelligence Review, vol. 59, no. 7, article 157
Supports: definition and categorisation of ZKML; main implementation bottlenecks
Locator: abstract; §III; Table VI
Version and catalogue details - S-1101 / Tier C
Proving LLMs at Scale ↗
Attestable · 2026 · Attestable blog
Supports: Attestable's reported prover, statement proven, performance and limits (provider-reported)
Version and catalogue details - S-1102 / Tier C
Pacing AI Requires Proof ↗
Attestable · 2026 · Attestable blog
Supports: Attestable's coverage argument and pacing proposal (provider-reported)
Version and catalogue details - S-1103 / Tier C
From Verifiability to Model-Weight Security ↗
Attestable · 2026 · Attestable blog
Supports: Attestable's proposal to prove randomly sampled outputs (provider-reported)
Version and catalogue details - S-0018 / Tier B
A System Overview for Near-Term, Low-Trust AI Compute Verification ↗
N. Cankaya · 2026 · Machine Intelligence Research Institute
Supports: ZKPs as a 'tentative plan B' in a verification system; overhead assessment; floating-point gap
Locator: §5.2.4
Version and catalogue details - S-1112 / Tier B
Hollow-LLM Attack: Computationally Trivial Weights in Zero-Knowledge Verification of LLM Inference ↗
C. Gong, B. Liu, M. Li · 2026 · arXiv
Supports: independent Hollow-LLM analysis: proofs do not bind computational effort; ghost-weight constructions; zkGPT-based experiment and cost results; countermeasures
Locator: Abstract; §I contributions; §V Table 2; §VI
Version and catalogue details - S-3060 / Tier A
zkGPT: An Efficient Non-interactive Zero-knowledge Proof Framework for LLM Inference ↗
W. Qu, Y. Sun, X. Liu, T. Lu, Y. Guo, K. Chen, J. Zhang · 2025 · 34th USENIX Security Symposium (USENIX Security 25), pp. 2045–2063
Supports: zkGPT design, non-interactive proofs, GPT-2 proving and verification figures, public code
Locator: abstract; §3; §6 Table 3
Version and catalogue details - S-1807 / Tier C
DeepProve-1: The First zkML System to Prove a Full LLM Inference ↗
Lagrange Labs · 2025 · Lagrange blog
Supports: Lagrange's reported proof of full GPT-2 inference and 'production-ready' description (provider-reported)
Version and catalogue details - S-1808 / Tier B
Lagrange-Labs/deep-prove (GitHub repository) ↗
Lagrange Labs · 2026 · GitHub
Supports: DeepProve public code, licence and reported GPT-2 proving figures (provider-reported)
Locator: README
Version and catalogue details
- Source review date
- 2026-09-25
- Drafted by (source map)
- ai
- Review handles (source map)
- codex-review