I-0007 / On-chip & hardware

Attestable Audits

A research prototype that runs AI safety benchmarks inside a trusted execution environment and publishes attestations binding the model, the audit and the results.

R2 DemonstratedSource reviewed 2026-09-25Provider-reported evidence

01 / The mechanism and its boundary

What the technique establishes

Attestable Audits, from University of Cambridge researchers, lets a model provider and an auditor who do not trust each other run a benchmark on a confidential model. Both send their inputs, encrypted, to a trusted execution environment: the provider its weights, the auditor its test code and data. The enclave runs the audit and publishes an attestation that binds the model hash, the audit and the result. Later, users can check that the model answering them is the one that was audited. The authors' prototype on CPU-only AWS Nitro Enclaves ran a 4-bit Llama-3.1-8B on MMLU, XSum and ToxicChat. CPU inference cost 21.7 times as much per token as GPU inference, and the enclave roughly doubled the CPU cost; the authors expect an overhead as small as 5 times on confidential-computing GPUs. No code is linked. The design trusts the TEE vendor, and prompt-based model exfiltration remains an open gap.

Threat model
Semi-trusted prover
Adversarial evaluation
Published analysis
Hardware needed
Existing hardware features
Prover cooperation
Required
Confidentiality
Preserving
Category
On chip & hardware
Technical detail and cited results
  • Security goals. G1 model verifiability, G2 audit verifiability, G3 confidentiality of model IP and audit data, G4 transparency of artifacts, G5 statelessness, and G6 output verifiability S-0009.
  • Adversaries. A1 is a network adversary that can intercept, tamper with or spoof messages, with denial of service excluded. A2 is a physical or privileged adversary able to take RAM snapshots, roll back VMs and run side-channel attacks S-0009.
  • PREPARE. The enclave generates a KEM key pair and attests its boot image. The provider sends the encrypted model, which the enclave quantizes and hashes. The enclave publishes an attestation linking model and quantized model to a transparency log S-0009.
  • ATTESTABLEAUDIT. The provider and the auditor send the encrypted model and the encrypted audit code and data. The enclave runs the audit in a sandbox and publishes an attestation binding model, audit and result S-0009.
  • INFERENCE. The enclave loads the earlier attestations, checks that the provider's model hash matches, and returns each encrypted response with an attestation linking model, prompt, output and audit result S-0009.
  • Verification. Verifiers check platform configuration registers (PCRs) against known images and check signatures with the TEE vendor's key or attestation service. Including the base image in the measurement allows revocation when vulnerabilities are found S-0009.
  • Prototype. The implementation is written in Rust with bindings to llama.cpp. A 4-core enclave on an m5.2xlarge instance is compared with CPU baselines on m5.xlarge and m5.2xlarge instances and an NVIDIA L40S GPU baseline. Enclave throughput was 1.84 tokens/s at $5.80 per 100K tokens, against 202 tokens/s and $0.12 on the GPU baseline. Zero-shot MMLU accuracy was 51.4% at 4-bit (57.4% excluding unparsable responses). The authors call this similar to the unquantized model on the GPU baseline, which scored 54.6% in the text and 58.9% in Table 2. Copying models into the enclave took at most 2 minutes S-0009.

Claims and scope

A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.

Readiness for a stated use

R2 Demonstrated

Assessed use: showing users that the model answering them is the audited one

low confidence · current · assessed 2026-09-25 · rubric 1.1

This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.

One developer paper reports end-to-end results on commercial hardware, but there is no public code and no independent reproduction, so confidence is low.

  • R1 met: the protocol, goals and adversaries are published S-0009.
  • R2 met through reproducible published end-to-end results. The paper specifies the protocol and its algorithms, the enclave and baseline instances, the models and the benchmarks. It reports results for Llama-3.1-8B on AWS Nitro Enclaves, which are commercially available hardware, against a stated adversary S-0009. The reported evaluation covers the audit step. The paper links no code, which the rubric does not require for this route.
  • R3 not met: no party other than the developers is documented relying on it for a verification decision, and the prototype is CPU-only research code rather than a production-grade system.
  • R4 not met: as of September 2026 no independent security evaluation has been published.

Evidence needed for the next level

  • Reliance by a party other than the developers for a verification decision, or a production-grade, available implementation, for example on GPU confidential computing at realistic model scale.

  • An independent public security evaluation (audit, red-team or peer-reviewed analysis).

Limitations, flaws, and blockers

These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.

significant / open / theoretical argument

Relies on the TEE vendor and inherits TEE attacks

The design depends on trusting the TEE vendor, AWS in the prototype S-0009. The authors cite memory-aliasing, ciphertext side-channel and malicious-interrupt attacks on confidential VMs (BadRAM, CIPHERLEAKS, Heckler). Their answer is to revoke vulnerable base images once such attacks are discovered S-0009.

S-0009
Response recorded by the source map

The authors propose revoking vulnerable base images; they do not report a red-team evaluation of the prototype S-0009.

significant / open / open question

Prompt-based model exfiltration is a residual gap

The authors state that "prompt-based model exfiltration during the user interaction step remains a residual gap" S-0009.

S-0009

minor / open / open question

CPU-only enclaves force small, quantized models and high cost

Memory limits required 4-bit quantization, and the quantized model scored 51.4% on zero-shot MMLU. CPU inference cost 21.7 times as much per token as GPU inference, and the enclave roughly doubled the CPU cost S-0009. The authors wrote that H100 confidential computing had no multi-GPU support S-0009. NVIDIA's white paper of August 2025 describes a protected-PCIe mode that passes all eight GPUs of a Hopper HGX node to one confidential VM, with NVLink traffic unencrypted S-1200.

S-0009S-1200

What still blocks use or stronger assurance

  1. The prototype needs porting to GPU confidential computing to handle larger models; the authors expect an overhead as small as 5 times there.

    Dependency: TEE remote attestation for AI workloads

    S-0009
  2. As of September 2026 no code has been released for the prototype.

    S-0009

Connections in the research map

Depends on

Mechanisms implemented

Concepts used

Organizations and developers

The Consortium’s case files

Related editorial reviews use the Consortium’s own descriptive scores and review dates. Their scores are separate from the atlas readiness rubric.

WV-06 / WorkloadInspect everything, reveal nothingRead case file ↗

Sources and provenance

  1. S-0009 / Tier B

    Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments ↗

    C. Schnabl, D. Hugenroth, B. Marino, A. R. Beresford · 2025 · ICML 2025 Workshop on Technical AI Governance

    Supports: all protocol, prototype, results and limitations statements

    Locator: §2 (background, GPU CC), §3 (goals, adversaries, Algorithms 1-3), §5 (Table 2), §7 (limitations)

    Version and catalogue details
  2. S-1200 / Tier B

    NVIDIA Secure AI with Blackwell and Hopper GPUs (White Paper) ↗

    NVIDIA · 2025 · NVIDIA documentation

    Supports: Hopper protected-PCIe multi-GPU mode with unencrypted NVLink

    Locator: p. 13

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review