01 / The mechanism and its boundary
What the technique establishes
Safeguard attestation aims to let users, auditors or other governments check that an AI service ran the safeguards it declares, such as a safety classifier, filter or usage monitor. Published designs run the safeguard code in a trusted execution environment, which signs a measurement of it with a hash of each input and response. A research prototype with public code does this on AWS Nitro Enclaves for a wrapper that routes an AI agent's traffic through an open-source guardrail. The guardrail model is called through an external API outside the enclave. In September 2026 Tinfoil reported rolling out enclave-run safeguard models with public, attested code. No independent security evaluation or reliance by another party has been documented. The main obstacles are showing that all traffic took the attested path and scaling to frontier GPU clusters. Attestation shows a safeguard ran, not that it works, since guardrails can be jailbroken.
- Threat model
- Semi-trusted prover
- Adversarial evaluation
- Published analysis
- Hardware needed
- Existing hardware features
- Prover cooperation
- Required
- Confidentiality
- Partial
- Category
- Cryptography & computation
Technical detail and cited results
- Proof-of-guardrail protocol. (1) A wrapper program f bundles the public guardrail g, its configuration and the mediation of all agent inputs and outputs. (2) When f is loaded, the enclave records a measurement m, a hash that depends on the binary of f. (3) The private agent is loaded afterwards as a secret input, so it is not part of m. (4) For a user input x and response r, f returns a document signed with the platform's attestation key that contains m and d = Hash(x, r). (5) The verifier checks the certificate chain against the platform's published root, compares m with the expected measurement of the open-source f, and checks d S-1500.
- Prototype costs. On AWS Nitro Enclaves, the authors report 34% added latency on average over the same agent and guardrail run outside the enclave (24.8–38.0% across the four measured steps), 97.8 ms to generate an attestation and 5.1 ms to verify it. Holding the whole guardrail runtime in enclave memory needs an m5.xlarge instance, which costs 18.5 times as much per hour as a t3.micro S-1500.
- PAL*M. It defines single and session inference properties, r = M(M_tok(q)) and its multi-turn form over the chat history. It binds hashes of the query, tokenizer, model and response into an Intel TDX quote, together with an NVIDIA H100 attestation token and a verifier challenge S-0012. PAL*M's added time was 3.8–11.4% of total run time for session inference and 45.5–66.4% for single prompts, across Llama-3.1-8B, Gemma-3-4B and Phi-4-Mini, so one attested prompt took 1.8 to 3 times as long as on the same TDX machine without PAL*M's measurements S-0012.
Claims and scope
A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.
Declared safeguards were applied during inference
Attests that a measured safeguard program (guardrail, filter, monitor) mediated the attested responses; coverage of all traffic is not established.
The declared model is the one being served
Property and audit attestations bind responses to a measured model (S-0012, S-0009).
Readiness for a stated use
Assessed use: attesting that a declared safeguard mediated a service's responses
low confidence · current · assessed 2026-09-25 · rubric 1.1
This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.
R2, narrowly: one prototype with public code attests a guardrail end to end on cloud enclaves. No independent security evaluation or reliance by another party has been documented.
- R1 met: designs that state the claim (a response was produced after a specific guardrail ran) and their trust assumptions are published S-1500, alongside designs for attesting inference properties S-0012 and plan-scoped monitoring S-1503.
- R2 met through the proof-of-guardrail prototype. Its code is public S-1501, and its authors report end-to-end results on production cloud enclave hardware (AWS Nitro Enclaves) against a stated adversary, a developer who skips or modifies the guardrail S-1500. PAL*M adds attested session inference on Intel TDX with an H100 GPU, but its properties do not cover safeguards S-0012.
- R3 not met: no party other than a developer is documented as relying on safeguard attestation, and the proof-of-guardrail code is described as a proof of concept that is not production-ready S-1501. Tinfoil reports that it is rolling out enclave-run safeguards in its chat service and can prove to an auditor or regulator that they are running, but as of mid-September 2026 the rollout was still under way S-3362.
- R4 not met: no independent audit or red-team has been published, and the monitoring prototype has not been stress-tested by a counterparty S-1504.
Confidence is low because the demonstration is far from a frontier serving stack. It reaches both its guardrail model and the agent's backend model through external APIs, attests only the responses for which attestation is offered, and runs on CPU enclaves S-1500. Its README states that the enclave does not yet restrict the agent's command execution, which could be used to bypass the guardrail S-1501. Intel TDX attestations have also been forged by attackers with physical access to the memory bus S-1202 S-3126.
Evidence needed for the next level
Reliance by a party other than the developer on safeguard attestations for a verification decision, or a production-grade system that is generally available.
An independent public evaluation (audit, red-team or peer-reviewed security analysis) of a safeguard-attestation system.
A demonstration in which the safeguard model itself runs inside the attested boundary on GPU hardware at a realistic serving scale.
A published way to show that all of a service's traffic, not only attested responses, passed through the attested safeguard path.
Limitations, flaws, and blockers
These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.
significant / open / theoretical argument
Attestation shows a safeguard ran, not that it is effective
Proof of guardrail ensures that the guardrail executed, but the guardrail can still err or be jailbroken. Because the guardrail must be open source, a malicious developer can attack it with jailbreaks while still presenting a valid proof. In the authors' evaluation, Llama Guard 3 reached an F1 score of 0.56 on the unsafe class of the ToxicChat dataset. The authors state that proof of guardrail should not be interpreted or advertised as proof of safety.
significant / open / theoretical argument
Selective attestation leaves traffic uncovered
Attestations are issued per response. In the prototype, the agent offers them when it receives high-stakes questions, so nothing shows that unattested traffic went through the same path. PAL*M's authors note that a prover could cherry-pick favourable executions, and suggest verifier-published nonces or requesting only session-level proofs. A governance analysis notes that auditors also need assurance that all activity is accounted for, since a host could start a second confidential virtual machine that bypasses monitoring.
significant / open / theoretical argument
Measurements may omit behaviour-relevant configuration or runtime changes
Every component that influences inference behaviour must be covered by the launch measurement, including feature flags, environment variables and invocation arguments. A launch measurement also does not show that a program keeps running as measured if the kernel is later compromised.
significant / open / theoretical argument
Components outside the attested boundary
In the proof-of-guardrail experiments, the guardrail model and the agent's backend model were both reached through external APIs, and the authors leave the decision to trust those APIs to the verifier. The measured wrapper must also have no vulnerability that lets the unmeasured agent bypass the guardrail, for example by executing arbitrary commands inside the enclave. The code's README states that the enclave does not currently restrict the agent's arbitrary command execution, which could be used to bypass guardrails.
significant / open / demonstrated attack
Memory-bus interposition extracts attestation keys and forges attestations
The TEE findings cover DDR5 attacks on Intel TDX, the H100 relay demonstration, DDR4 attacks on AMD SEV-SNP, and software-only SEV-SNP forgery before AMD's fixes S-1202 S-3126 S-1210 S-1212 S-1213. These are inherited hardware limits; a governance analysis explains why physical access matters in a treaty setting S-0014.
What still blocks use or stronger assurance
- S-1500S-0014
No published design shows that all of a provider's traffic passes through the attested safeguard path; current evidence covers individual attested responses.
Frontier model inference typically needs several GPUs, GPU confidential computing is less mature than CPU support, and CPU inference, which an enclave prototype had to use, ran about 100 times slower than GPU inference.
Dependency: TEE remote attestation for AI workloads
S-0014S-0009Trust rests on a small number of hardware vendors, and a per-CPU Intel attestation key has been extracted by physical attack.
Dependency: TEE remote attestation for AI workloads
S-0014S-1202Safeguard evidence must be bound to the model actually served, which depends on model-identity attestation.
Dependency: Model identity attestation
S-0009S-0013- S-1501S-1504
No independent red-team or audit of a safeguard-attestation system has been published, and the available prototypes are described by their authors as proofs of concept that have not been stress-tested by a counterparty.
Connections in the research map
Depends on
- TEE remote attestation for AI workloads
Current designs rely on TEE measurement and remote attestation.
- Model identity attestation
Safeguard evidence is meaningful only when bound to the model actually served.
Complementary techniques
Alternative approaches
Concepts used
Organizations and developers
The Consortium’s case files
Related editorial reviews use the Consortium’s own descriptive scores and review dates. Their scores are separate from the atlas readiness rubric.
RT-09 / AttestationTamperproof certificates, continuouslyRead case file ↗Sources and provenance
- S-1500 / Tier B
Proof-of-Guardrail in AI Agents and What (Not) to Trust from It ↗
X. Jin, M. Duan, Q. Lin, A. Chan, Z. Chen, J. Du, X. Ren · 2026 · arXiv
Supports: problem statement; protocol; threat model and trust assumptions; implementation and overheads; tamper tests; guardrail accuracy on ToxicChat; jailbreak risk; proof-of-safety caveat; external APIs; wrapper-bypass risk; selective attestation
Locator: abstract; §3; §4.1; Tables 1-3; Appendix A
Version and catalogue details - S-1501 / Tier B
Verifiable-ClawGuard: proof-of-guardrail reference code ↗
SaharaLabsAI · 2026 · GitHub
Supports: public code; proof-of-concept status; stated limitation on agent command execution
Locator: README, including Limitations
Version and catalogue details - S-3362 / Tier C
Safety Without Compromising on Privacy ↗
D. McCann-Sayles, S. Servan-Schreiber, T. Verma · 2026 · Tinfoil blog
Supports: Tinfoil's enclave-run safeguard pipeline: models, what leaves the enclave, open-source and attested code, rollout status
Locator: whole post (updated 2026-09-16)
Version and catalogue details - S-0012 / Tier B
PAL*M: Property Attestation for Large Generative Models ↗
P. Chantasantitam, A. I. Caulfield, V. Duddu, L. J. Gunn, N. Asokan · 2026 · arXiv
Supports: inference property definitions; TDX+H100 implementation; overheads and baseline; threat model; cherry-picking discussion
Locator: abstract; §3.2; §4.3.4 (Defs. 7-8); §4.4; Table 6; Appendix A
Version and catalogue details - S-0009 / Tier B
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments ↗
C. Schnabl, D. Hugenroth, B. Marino, A. R. Beresford · 2025 · ICML 2025 Workshop on Technical AI Governance
Supports: inference protocol linking model, audit result, prompt and response; CPU-only enclave prototype; CPU versus GPU cost and slowdown
Locator: §3 (Inference protocol); §5; Table 2
Version and catalogue details - S-1503 / Tier B
Enabling Verifiably-Scoped Monitoring through Large Language Models and Trusted Compute ↗
B. Penchas, G. Zhao, R. Rinberg · 2026 · ICML 2026 Workshop on Technical AI Governance Research
Supports: verifiably-scoped monitoring protocol
Locator: abstract
Version and catalogue details - S-1504 / Tier C
Auditor-in-a-Box: Tools for Third-Party Auditing ↗
R. Rinberg, B. Penchas · 2026 · LessWrong
Supports: reference implementation in Tinfoil confidential VMs; stated limitations
Locator: reference implementation; limitations
Version and catalogue details - S-0014 / Tier C
On TEEs for Privacy-Preserving Monitoring in AI Governance ↗
Gloria Z · 2026 · MIRI Technical Governance Team
Supports: policy adherence as a verification property; measurement completeness; runtime state; second-CVM completeness gap; vendor root of trust; treaty threat model; GPU TEE maturity and multi-GPU inference
Locator: deployment integrity; hardware auditability; resource accounting; physical attack surface
Version and catalogue details - S-0003 / Tier B
Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies ↗
M. Brundage, N. Dreksler, A. Homewood, S. McGregor, P. Paskov, C. Stosz, G. Sastry, A. F. Cooper, G. Balston, S. Adler, S. Casper, M. Anderljung, G. Werner, S. Mindermann, V. Mavroudis, B. Bucknall, C. Stix, J. Freund, L. Pacchiardi, J. Hernandez-Orallo, M. Pistillo, M. Chen, C. Painter, D. W. Ball, C. O'Keefe, G. Weil, B. Harack, G. Finley, R. Hassan, S. Emmons, C. Foster, A. Reuel, B. Treece, Y. Bengio, D. Reti, R. Bommasani, C. Trout, A. S. Shamsabadi, R. Dattani, A. Weller, R. Trager, J. Sevilla, L. Wagner, L. Soder, K. Ramakrishnan, H. Papadatos, M. Murray, R. Tovcimak · 2026 · arXiv
Supports: configuration drift, such as swapping safety classifiers or relaxing filter thresholds
Locator: §5.2
Version and catalogue details - S-0018 / Tier B
A System Overview for Near-Term, Low-Trust AI Compute Verification ↗
N. Cankaya · 2026 · Machine Intelligence Research Institute
Supports: verifier-side compliance screening of re-executed records
Locator: §3.2.2; §5.2.3
Version and catalogue details - S-1502 / Tier B
Has My System Prompt Been Used? Large Language Model Prompt Membership Inference ↗
R. Levin, V. Cherepanova, A. Hans, A. Schwarzschild, T. Goldstein · 2025 · arXiv
Supports: black-box statistical test for system-prompt use; its prompt-protection setting
Locator: abstract; §3.2
Version and catalogue details - S-1202 / Tier A
TEE.fail: Breaking Trusted Execution Environments via DDR5 Memory Bus Interposition ↗
J. Chuang, A. Seto, N. Berrios, S. van Schaik, C. Garman, D. Genkin · 2026 · 2026 IEEE Symposium on Security and Privacy (SP)
Supports: physical key extraction from TDX and signing-key extraction in SEV-SNP; forged attestations against NVIDIA GPU confidential computing; cost; vendor acknowledgement and positions
Locator: project site summary; paper abstract and disclosure
Version and catalogue details - S-3126 / Tier A
DDRop: Active Memory Interposer Attacks on Confidential VMs by Dropping DDR5 Writes ↗
J. De Meulemeester, S. Gloor, P. Jattke, D. Moghimi, D. Oswald, M. Thompson, K. Razavi, I. Verbauwhede, J. Van Bulck · 2026 · 2026 ACM SIGSAC Conference on Computer and Communications Security (CCS '26)
Supports: DDRop forges attestation reports on an up-to-date Intel TDX platform with an active DDR5 interposer
Locator: site summary
Version and catalogue details - S-1210 / Tier A
Battering RAM: Low-Cost Interposer Attacks on Confidential Computing via Dynamic Memory Aliasing ↗
J. De Meulemeester, D. Oswald, I. Verbauwhede, J. Van Bulck · 2026 · 47th IEEE Symposium on Security and Privacy (S&P 2026)
Supports: Battering RAM forges SEV-SNP attestation with a DDR4 interposer
Version and catalogue details - S-1212 / Tier A
RMPocalypse: How a Catch-22 Breaks AMD SEV-SNP ↗
B. Schlüter, S. Shinde · 2025 · 2025 ACM SIGSAC Conference on Computer and Communications Security (CCS '25)
Supports: RMPocalypse forges SEV-SNP attestation from a malicious hypervisor
Version and catalogue details - S-1213 / Tier B
SEV-SNP RMP Initialization Vulnerability (AMD-SB-3020) ↗
AMD · 2025 · AMD product security bulletin
Supports: AMD firmware fixes for RMPocalypse (CVE-2025-0033)
Version and catalogue details - S-0013 / Tier C
How Tinfoil Proves Exactly What Model Is Running ↗
Tinfoil Team · 2026 · Tinfoil
Supports: binding model weights to enclave attestation
Locator: whole post
Version and catalogue details
- Source review date
- 2026-09-25
- Drafted by (source map)
- ai
- Review handles (source map)
- codex-review