01 / The mechanism and its boundary
What the technique establishes
Attestable Audits, from University of Cambridge researchers, lets a model provider and an auditor who do not trust each other run a benchmark on a confidential model. Both send their inputs, encrypted, to a trusted execution environment: the provider its weights, the auditor its test code and data. The enclave runs the audit and publishes an attestation that binds the model hash, the audit and the result. Later, users can check that the model answering them is the one that was audited. The authors' prototype on CPU-only AWS Nitro Enclaves ran a 4-bit Llama-3.1-8B on MMLU, XSum and ToxicChat. CPU inference cost 21.7 times as much per token as GPU inference, and the enclave roughly doubled the CPU cost; the authors expect an overhead as small as 5 times on confidential-computing GPUs. No code is linked. The design trusts the TEE vendor, and prompt-based model exfiltration remains an open gap.
- Threat model
- Semi-trusted prover
- Adversarial evaluation
- Published analysis
- Hardware needed
- Existing hardware features
- Prover cooperation
- Required
- Confidentiality
- Preserving
- Category
- On chip & hardware
Technical detail and cited results
- Security goals. G1 model verifiability, G2 audit verifiability, G3 confidentiality of model IP and audit data, G4 transparency of artifacts, G5 statelessness, and G6 output verifiability S-0009.
- Adversaries. A1 is a network adversary that can intercept, tamper with or spoof messages, with denial of service excluded. A2 is a physical or privileged adversary able to take RAM snapshots, roll back VMs and run side-channel attacks S-0009.
- PREPARE. The enclave generates a KEM key pair and attests its boot image. The provider sends the encrypted model, which the enclave quantizes and hashes. The enclave publishes an attestation linking model and quantized model to a transparency log S-0009.
- ATTESTABLEAUDIT. The provider and the auditor send the encrypted model and the encrypted audit code and data. The enclave runs the audit in a sandbox and publishes an attestation binding model, audit and result S-0009.
- INFERENCE. The enclave loads the earlier attestations, checks that the provider's model hash matches, and returns each encrypted response with an attestation linking model, prompt, output and audit result S-0009.
- Verification. Verifiers check platform configuration registers (PCRs) against known images and check signatures with the TEE vendor's key or attestation service. Including the base image in the measurement allows revocation when vulnerabilities are found S-0009.
- Prototype. The implementation is written in Rust with bindings to llama.cpp. A 4-core enclave on an m5.2xlarge instance is compared with CPU baselines on m5.xlarge and m5.2xlarge instances and an NVIDIA L40S GPU baseline. Enclave throughput was 1.84 tokens/s at $5.80 per 100K tokens, against 202 tokens/s and $0.12 on the GPU baseline. Zero-shot MMLU accuracy was 51.4% at 4-bit (57.4% excluding unparsable responses). The authors call this similar to the unquantized model on the GPU baseline, which scored 54.6% in the text and 58.9% in Table 2. Copying models into the enclave took at most 2 minutes S-0009.
Claims and scope
A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.
The declared model is the one being served
Users can check that the model answering them is the audited one.
The declared evaluation was run (upstream draft)
The audit protocol binds the model hash, audit code and data, and result in a published attestation (S-0009).
Readiness for a stated use
Assessed use: showing users that the model answering them is the audited one
low confidence · current · assessed 2026-09-25 · rubric 1.1
This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.
One developer paper reports end-to-end results on commercial hardware, but there is no public code and no independent reproduction, so confidence is low.
- R1 met: the protocol, goals and adversaries are published S-0009.
- R2 met through reproducible published end-to-end results. The paper specifies the protocol and its algorithms, the enclave and baseline instances, the models and the benchmarks. It reports results for Llama-3.1-8B on AWS Nitro Enclaves, which are commercially available hardware, against a stated adversary S-0009. The reported evaluation covers the audit step. The paper links no code, which the rubric does not require for this route.
- R3 not met: no party other than the developers is documented relying on it for a verification decision, and the prototype is CPU-only research code rather than a production-grade system.
- R4 not met: as of September 2026 no independent security evaluation has been published.
Evidence needed for the next level
Reliance by a party other than the developers for a verification decision, or a production-grade, available implementation, for example on GPU confidential computing at realistic model scale.
An independent public security evaluation (audit, red-team or peer-reviewed analysis).
Limitations, flaws, and blockers
These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.
significant / open / theoretical argument
Relies on the TEE vendor and inherits TEE attacks
The design depends on trusting the TEE vendor, AWS in the prototype S-0009. The authors cite memory-aliasing, ciphertext side-channel and malicious-interrupt attacks on confidential VMs (BadRAM, CIPHERLEAKS, Heckler). Their answer is to revoke vulnerable base images once such attacks are discovered S-0009.
The authors propose revoking vulnerable base images; they do not report a red-team evaluation of the prototype S-0009.
significant / open / open question
Prompt-based model exfiltration is a residual gap
The authors state that "prompt-based model exfiltration during the user interaction step remains a residual gap" S-0009.
minor / open / open question
CPU-only enclaves force small, quantized models and high cost
Memory limits required 4-bit quantization, and the quantized model scored 51.4% on zero-shot MMLU. CPU inference cost 21.7 times as much per token as GPU inference, and the enclave roughly doubled the CPU cost S-0009. The authors wrote that H100 confidential computing had no multi-GPU support S-0009. NVIDIA's white paper of August 2025 describes a protected-PCIe mode that passes all eight GPUs of a Hopper HGX node to one confidential VM, with NVLink traffic unencrypted S-1200.
What still blocks use or stronger assurance
The prototype needs porting to GPU confidential computing to handle larger models; the authors expect an overhead as small as 5 times there.
Dependency: TEE remote attestation for AI workloads
S-0009- S-0009
As of September 2026 no code has been released for the prototype.
Connections in the research map
Depends on
- TEE remote attestation for AI workloads
Built on AWS Nitro Enclaves attestation.
Mechanisms implemented
Concepts used
Organizations and developers
The Consortium’s case files
Related editorial reviews use the Consortium’s own descriptive scores and review dates. Their scores are separate from the atlas readiness rubric.
WV-06 / WorkloadInspect everything, reveal nothingRead case file ↗Sources and provenance
- S-0009 / Tier B
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments ↗
C. Schnabl, D. Hugenroth, B. Marino, A. R. Beresford · 2025 · ICML 2025 Workshop on Technical AI Governance
Supports: all protocol, prototype, results and limitations statements
Locator: §2 (background, GPU CC), §3 (goals, adversaries, Algorithms 1-3), §5 (Table 2), §7 (limitations)
Version and catalogue details - S-1200 / Tier B
NVIDIA Secure AI with Blackwell and Hopper GPUs (White Paper) ↗
NVIDIA · 2025 · NVIDIA documentation
Supports: Hopper protected-PCIe multi-GPU mode with unencrypted NVLink
Locator: p. 13
Version and catalogue details
- Source review date
- 2026-09-25
- Drafted by (source map)
- ai
- Review handles (source map)
- codex-review