M-0023 / Cryptography & computation

Safeguard attestation

Hardware-signed evidence that an AI service ran its declared safeguards, such as a guardrail classifier or monitor, when producing a given response.

R2 DemonstratedSource reviewed 2026-09-25Provider-reported evidence

01 / The mechanism and its boundary

What the technique establishes

Safeguard attestation aims to let users, auditors or other governments check that an AI service ran the safeguards it declares, such as a safety classifier, filter or usage monitor. Published designs run the safeguard code in a trusted execution environment, which signs a measurement of it with a hash of each input and response. A research prototype with public code does this on AWS Nitro Enclaves for a wrapper that routes an AI agent's traffic through an open-source guardrail. The guardrail model is called through an external API outside the enclave. In September 2026 Tinfoil reported rolling out enclave-run safeguard models with public, attested code. No independent security evaluation or reliance by another party has been documented. The main obstacles are showing that all traffic took the attested path and scaling to frontier GPU clusters. Attestation shows a safeguard ran, not that it works, since guardrails can be jailbroken.

Threat model
Semi-trusted prover
Adversarial evaluation
Published analysis
Hardware needed
Existing hardware features
Prover cooperation
Required
Confidentiality
Partial
Category
Cryptography & computation
Technical detail and cited results
  • Proof-of-guardrail protocol. (1) A wrapper program f bundles the public guardrail g, its configuration and the mediation of all agent inputs and outputs. (2) When f is loaded, the enclave records a measurement m, a hash that depends on the binary of f. (3) The private agent is loaded afterwards as a secret input, so it is not part of m. (4) For a user input x and response r, f returns a document signed with the platform's attestation key that contains m and d = Hash(x, r). (5) The verifier checks the certificate chain against the platform's published root, compares m with the expected measurement of the open-source f, and checks d S-1500.
  • Prototype costs. On AWS Nitro Enclaves, the authors report 34% added latency on average over the same agent and guardrail run outside the enclave (24.8–38.0% across the four measured steps), 97.8 ms to generate an attestation and 5.1 ms to verify it. Holding the whole guardrail runtime in enclave memory needs an m5.xlarge instance, which costs 18.5 times as much per hour as a t3.micro S-1500.
  • PAL*M. It defines single and session inference properties, r = M(M_tok(q)) and its multi-turn form over the chat history. It binds hashes of the query, tokenizer, model and response into an Intel TDX quote, together with an NVIDIA H100 attestation token and a verifier challenge S-0012. PAL*M's added time was 3.8–11.4% of total run time for session inference and 45.5–66.4% for single prompts, across Llama-3.1-8B, Gemma-3-4B and Phi-4-Mini, so one attested prompt took 1.8 to 3 times as long as on the same TDX machine without PAL*M's measurements S-0012.

Claims and scope

A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.

Readiness for a stated use

R2 Demonstrated

Assessed use: attesting that a declared safeguard mediated a service's responses

low confidence · current · assessed 2026-09-25 · rubric 1.1

This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.

R2, narrowly: one prototype with public code attests a guardrail end to end on cloud enclaves. No independent security evaluation or reliance by another party has been documented.

  • R1 met: designs that state the claim (a response was produced after a specific guardrail ran) and their trust assumptions are published S-1500, alongside designs for attesting inference properties S-0012 and plan-scoped monitoring S-1503.
  • R2 met through the proof-of-guardrail prototype. Its code is public S-1501, and its authors report end-to-end results on production cloud enclave hardware (AWS Nitro Enclaves) against a stated adversary, a developer who skips or modifies the guardrail S-1500. PAL*M adds attested session inference on Intel TDX with an H100 GPU, but its properties do not cover safeguards S-0012.
  • R3 not met: no party other than a developer is documented as relying on safeguard attestation, and the proof-of-guardrail code is described as a proof of concept that is not production-ready S-1501. Tinfoil reports that it is rolling out enclave-run safeguards in its chat service and can prove to an auditor or regulator that they are running, but as of mid-September 2026 the rollout was still under way S-3362.
  • R4 not met: no independent audit or red-team has been published, and the monitoring prototype has not been stress-tested by a counterparty S-1504.

Confidence is low because the demonstration is far from a frontier serving stack. It reaches both its guardrail model and the agent's backend model through external APIs, attests only the responses for which attestation is offered, and runs on CPU enclaves S-1500. Its README states that the enclave does not yet restrict the agent's command execution, which could be used to bypass the guardrail S-1501. Intel TDX attestations have also been forged by attackers with physical access to the memory bus S-1202 S-3126.

Evidence needed for the next level

  • Reliance by a party other than the developer on safeguard attestations for a verification decision, or a production-grade system that is generally available.

  • An independent public evaluation (audit, red-team or peer-reviewed security analysis) of a safeguard-attestation system.

  • A demonstration in which the safeguard model itself runs inside the attested boundary on GPU hardware at a realistic serving scale.

  • A published way to show that all of a service's traffic, not only attested responses, passed through the attested safeguard path.

Limitations, flaws, and blockers

These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.

significant / open / theoretical argument

Attestation shows a safeguard ran, not that it is effective

Proof of guardrail ensures that the guardrail executed, but the guardrail can still err or be jailbroken. Because the guardrail must be open source, a malicious developer can attack it with jailbreaks while still presenting a valid proof. In the authors' evaluation, Llama Guard 3 reached an F1 score of 0.56 on the unsafe class of the ToxicChat dataset. The authors state that proof of guardrail should not be interpreted or advertised as proof of safety.

S-1500

significant / open / theoretical argument

Selective attestation leaves traffic uncovered

Attestations are issued per response. In the prototype, the agent offers them when it receives high-stakes questions, so nothing shows that unattested traffic went through the same path. PAL*M's authors note that a prover could cherry-pick favourable executions, and suggest verifier-published nonces or requesting only session-level proofs. A governance analysis notes that auditors also need assurance that all activity is accounted for, since a host could start a second confidential virtual machine that bypasses monitoring.

S-1500S-0012S-0014

significant / open / theoretical argument

Measurements may omit behaviour-relevant configuration or runtime changes

Every component that influences inference behaviour must be covered by the launch measurement, including feature flags, environment variables and invocation arguments. A launch measurement also does not show that a program keeps running as measured if the kernel is later compromised.

S-0014

significant / open / theoretical argument

Components outside the attested boundary

In the proof-of-guardrail experiments, the guardrail model and the agent's backend model were both reached through external APIs, and the authors leave the decision to trust those APIs to the verifier. The measured wrapper must also have no vulnerability that lets the unmeasured agent bypass the guardrail, for example by executing arbitrary commands inside the enclave. The code's README states that the enclave does not currently restrict the agent's arbitrary command execution, which could be used to bypass guardrails.

S-1500S-1501

significant / open / demonstrated attack

Memory-bus interposition extracts attestation keys and forges attestations

The TEE findings cover DDR5 attacks on Intel TDX, the H100 relay demonstration, DDR4 attacks on AMD SEV-SNP, and software-only SEV-SNP forgery before AMD's fixes S-1202 S-3126 S-1210 S-1212 S-1213. These are inherited hardware limits; a governance analysis explains why physical access matters in a treaty setting S-0014.

S-1202S-3126S-1210S-1212S-1213S-0012S-0014S-1500S-0018
Response recorded by the source map

Intel and AMD place the physical attack class outside their threat models, according to the researchers. AMD reports firmware fixes for RMPocalypse S-1202 S-3126 S-1213.

What still blocks use or stronger assurance

  1. No published design shows that all of a provider's traffic passes through the attested safeguard path; current evidence covers individual attested responses.

    S-1500S-0014
  2. Frontier model inference typically needs several GPUs, GPU confidential computing is less mature than CPU support, and CPU inference, which an enclave prototype had to use, ran about 100 times slower than GPU inference.

    Dependency: TEE remote attestation for AI workloads

    S-0014S-0009
  3. Trust rests on a small number of hardware vendors, and a per-CPU Intel attestation key has been extracted by physical attack.

    Dependency: TEE remote attestation for AI workloads

    S-0014S-1202
  4. Safeguard evidence must be bound to the model actually served, which depends on model-identity attestation.

    Dependency: Model identity attestation

    S-0009S-0013
  5. No independent red-team or audit of a safeguard-attestation system has been published, and the available prototypes are described by their authors as proofs of concept that have not been stress-tested by a counterparty.

    S-1501S-1504

Connections in the research map

Depends on

Complementary techniques

Alternative approaches

Concepts used

Organizations and developers

The Consortium’s case files

Related editorial reviews use the Consortium’s own descriptive scores and review dates. Their scores are separate from the atlas readiness rubric.

RT-09 / AttestationTamperproof certificates, continuouslyRead case file ↗

Sources and provenance

  1. S-1500 / Tier B

    Proof-of-Guardrail in AI Agents and What (Not) to Trust from It ↗

    X. Jin, M. Duan, Q. Lin, A. Chan, Z. Chen, J. Du, X. Ren · 2026 · arXiv

    Supports: problem statement; protocol; threat model and trust assumptions; implementation and overheads; tamper tests; guardrail accuracy on ToxicChat; jailbreak risk; proof-of-safety caveat; external APIs; wrapper-bypass risk; selective attestation

    Locator: abstract; §3; §4.1; Tables 1-3; Appendix A

    Version and catalogue details
  2. S-1501 / Tier B

    Verifiable-ClawGuard: proof-of-guardrail reference code ↗

    SaharaLabsAI · 2026 · GitHub

    Supports: public code; proof-of-concept status; stated limitation on agent command execution

    Locator: README, including Limitations

    Version and catalogue details
  3. S-3362 / Tier C

    Safety Without Compromising on Privacy ↗

    D. McCann-Sayles, S. Servan-Schreiber, T. Verma · 2026 · Tinfoil blog

    Supports: Tinfoil's enclave-run safeguard pipeline: models, what leaves the enclave, open-source and attested code, rollout status

    Locator: whole post (updated 2026-09-16)

    Version and catalogue details
  4. S-0012 / Tier B

    PAL*M: Property Attestation for Large Generative Models ↗

    P. Chantasantitam, A. I. Caulfield, V. Duddu, L. J. Gunn, N. Asokan · 2026 · arXiv

    Supports: inference property definitions; TDX+H100 implementation; overheads and baseline; threat model; cherry-picking discussion

    Locator: abstract; §3.2; §4.3.4 (Defs. 7-8); §4.4; Table 6; Appendix A

    Version and catalogue details
  5. S-0009 / Tier B

    Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments ↗

    C. Schnabl, D. Hugenroth, B. Marino, A. R. Beresford · 2025 · ICML 2025 Workshop on Technical AI Governance

    Supports: inference protocol linking model, audit result, prompt and response; CPU-only enclave prototype; CPU versus GPU cost and slowdown

    Locator: §3 (Inference protocol); §5; Table 2

    Version and catalogue details
  6. S-1503 / Tier B

    Enabling Verifiably-Scoped Monitoring through Large Language Models and Trusted Compute ↗

    B. Penchas, G. Zhao, R. Rinberg · 2026 · ICML 2026 Workshop on Technical AI Governance Research

    Supports: verifiably-scoped monitoring protocol

    Locator: abstract

    Version and catalogue details
  7. S-1504 / Tier C

    Auditor-in-a-Box: Tools for Third-Party Auditing ↗

    R. Rinberg, B. Penchas · 2026 · LessWrong

    Supports: reference implementation in Tinfoil confidential VMs; stated limitations

    Locator: reference implementation; limitations

    Version and catalogue details
  8. S-0014 / Tier C

    On TEEs for Privacy-Preserving Monitoring in AI Governance ↗

    Gloria Z · 2026 · MIRI Technical Governance Team

    Supports: policy adherence as a verification property; measurement completeness; runtime state; second-CVM completeness gap; vendor root of trust; treaty threat model; GPU TEE maturity and multi-GPU inference

    Locator: deployment integrity; hardware auditability; resource accounting; physical attack surface

    Version and catalogue details
  9. S-0003 / Tier B

    Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies ↗

    M. Brundage, N. Dreksler, A. Homewood, S. McGregor, P. Paskov, C. Stosz, G. Sastry, A. F. Cooper, G. Balston, S. Adler, S. Casper, M. Anderljung, G. Werner, S. Mindermann, V. Mavroudis, B. Bucknall, C. Stix, J. Freund, L. Pacchiardi, J. Hernandez-Orallo, M. Pistillo, M. Chen, C. Painter, D. W. Ball, C. O'Keefe, G. Weil, B. Harack, G. Finley, R. Hassan, S. Emmons, C. Foster, A. Reuel, B. Treece, Y. Bengio, D. Reti, R. Bommasani, C. Trout, A. S. Shamsabadi, R. Dattani, A. Weller, R. Trager, J. Sevilla, L. Wagner, L. Soder, K. Ramakrishnan, H. Papadatos, M. Murray, R. Tovcimak · 2026 · arXiv

    Supports: configuration drift, such as swapping safety classifiers or relaxing filter thresholds

    Locator: §5.2

    Version and catalogue details
  10. S-0018 / Tier B

    A System Overview for Near-Term, Low-Trust AI Compute Verification ↗

    N. Cankaya · 2026 · Machine Intelligence Research Institute

    Supports: verifier-side compliance screening of re-executed records

    Locator: §3.2.2; §5.2.3

    Version and catalogue details
  11. S-1502 / Tier B

    Has My System Prompt Been Used? Large Language Model Prompt Membership Inference ↗

    R. Levin, V. Cherepanova, A. Hans, A. Schwarzschild, T. Goldstein · 2025 · arXiv

    Supports: black-box statistical test for system-prompt use; its prompt-protection setting

    Locator: abstract; §3.2

    Version and catalogue details
  12. S-1202 / Tier A

    TEE.fail: Breaking Trusted Execution Environments via DDR5 Memory Bus Interposition ↗

    J. Chuang, A. Seto, N. Berrios, S. van Schaik, C. Garman, D. Genkin · 2026 · 2026 IEEE Symposium on Security and Privacy (SP)

    Supports: physical key extraction from TDX and signing-key extraction in SEV-SNP; forged attestations against NVIDIA GPU confidential computing; cost; vendor acknowledgement and positions

    Locator: project site summary; paper abstract and disclosure

    Version and catalogue details
  13. S-3126 / Tier A

    DDRop: Active Memory Interposer Attacks on Confidential VMs by Dropping DDR5 Writes ↗

    J. De Meulemeester, S. Gloor, P. Jattke, D. Moghimi, D. Oswald, M. Thompson, K. Razavi, I. Verbauwhede, J. Van Bulck · 2026 · 2026 ACM SIGSAC Conference on Computer and Communications Security (CCS '26)

    Supports: DDRop forges attestation reports on an up-to-date Intel TDX platform with an active DDR5 interposer

    Locator: site summary

    Version and catalogue details
  14. S-1210 / Tier A

    Battering RAM: Low-Cost Interposer Attacks on Confidential Computing via Dynamic Memory Aliasing ↗

    J. De Meulemeester, D. Oswald, I. Verbauwhede, J. Van Bulck · 2026 · 47th IEEE Symposium on Security and Privacy (S&P 2026)

    Supports: Battering RAM forges SEV-SNP attestation with a DDR4 interposer

    Version and catalogue details
  15. S-1212 / Tier A

    RMPocalypse: How a Catch-22 Breaks AMD SEV-SNP ↗

    B. Schlüter, S. Shinde · 2025 · 2025 ACM SIGSAC Conference on Computer and Communications Security (CCS '25)

    Supports: RMPocalypse forges SEV-SNP attestation from a malicious hypervisor

    Version and catalogue details
  16. S-1213 / Tier B

    SEV-SNP RMP Initialization Vulnerability (AMD-SB-3020) ↗

    AMD · 2025 · AMD product security bulletin

    Supports: AMD firmware fixes for RMPocalypse (CVE-2025-0033)

    Version and catalogue details
  17. S-0013 / Tier C

    How Tinfoil Proves Exactly What Model Is Running ↗

    Tinfoil Team · 2026 · Tinfoil

    Supports: binding model weights to enclave attestation

    Locator: whole post

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review