01 / The mechanism and its boundary
What is being described
The claim is that specified safety measures, such as input filters, output checks or monitoring for misuse, actually ran on the requests a deployed model served. Rules for deploying AI systems often require such safeguards. A developer's statement that safeguards exist does not show that they ran on every request, or that the version checked is the one in production. Verifying application matters to regulators and to agreements that allow deployment on condition of mitigations. It is a positive claim, but it inherits the difficulties of verifying the served model and adds more. Safeguards are often separate software components whose configuration can change, and the verifier usually cannot see users' prompts or the provider's systems. Proposals use trusted execution environments to attest the serving stack, audits run inside such environments, and inspections. Whether a safeguard is effective is a separate question from whether it was applied.
State of verification
Editorial synthesis from the AI Verification Tech Map.
Safeguard application has been shown only in small research prototypes, which can show that a declared safeguard ran, not that it works.
Safeguard attestation (R2) runs the safeguard inside a trusted execution environment, whose hardware signs a measurement of its code with a commitment to each input and response S-1500. It builds on TEE remote attestation (R3) and on knowing which model is served (The declared model is the one being served). The proof-of-guardrail prototype demonstrates it on CPU enclaves in AWS and calls its guardrail model through an external API S-1500. In the authors' tests it detected modified guardrail code, attestations and responses S-1500. Confidential multi-party verification (R2) can limit monitoring to a plan both parties sign S-1503, and Auditor-in-a-Box demonstrates such plan-scoped monitoring, though its authors state that the demo's user data and plan execution are not actually secure S-1504.
Every component that influences inference must be covered by the launch measurement S-0014, and the prototype attests only the responses for which it offers attestation, so coverage of all traffic is not shown S-1500. Where hardware trust is unavailable, the sources fall back on inspections, audits and personnel-based layers S-0062 S-0003 S-0002.
Connections in the research map
Techniques addressing this claim
Sources and provenance
- S-0002 / Tier B
Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment ↗
M. Baker, G. Kulp, O. Marks, M. Brundage, L. Heim · 2025 · RAND Corporation
Supports: deployment mitigations specified by inputs and outputs (filters, oversight checks); difficulty of choosing technical rules; whistleblower and interview layers
Locator: Table 3; §4
Version and catalogue details - S-0062 / Tier B
Verification methods for international AI agreements ↗
A. R. Wasil, T. Reed, J. W. Miller, P. Barnett · 2024 · arXiv
Supports: AI developer inspections to check authorised code and implementation of evaluations and safeguards; software can be quickly modified or hidden
Locator: Access-dependent methods; Table 1
Version and catalogue details - S-0003 / Tier B
Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies ↗
M. Brundage, N. Dreksler, A. Homewood, S. McGregor, P. Paskov, C. Stosz, G. Sastry, A. F. Cooper, G. Balston, S. Adler, S. Casper, M. Anderljung, G. Werner, S. Mindermann, V. Mavroudis, B. Bucknall, C. Stix, J. Freund, L. Pacchiardi, J. Hernandez-Orallo, M. Pistillo, M. Chen, C. Painter, D. W. Ball, C. O'Keefe, G. Weil, B. Harack, G. Finley, R. Hassan, S. Emmons, C. Foster, A. Reuel, B. Treece, Y. Bengio, D. Reti, R. Bommasani, C. Trout, A. S. Shamsabadi, R. Dattani, A. Weller, R. Trager, J. Sevilla, L. Wagner, L. Soder, K. Ramakrishnan, H. Papadatos, M. Murray, R. Tovcimak · 2026 · arXiv
Supports: third-party verification of developers' safety and security claims with deep, secure access; AI Assurance Levels
Locator: abstract
Version and catalogue details - S-0014 / Tier C
On TEEs for Privacy-Preserving Monitoring in AI Governance ↗
Gloria Z · 2026 · MIRI Technical Governance Team
Supports: TEEs for policy-adherent inference; launch measurement must cover all components; attestation-key holder can forge reports
Locator: main argument; Limitations
Version and catalogue details - S-0009 / Tier B
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments ↗
C. Schnabl, D. Hugenroth, B. Marino, A. R. Beresford · 2025 · ICML 2025 Workshop on Technical AI Governance
Supports: TEE-based verifiable benchmarks keeping model IP and datasets confidential; prototype with Llama-3.1
Locator: abstract
Version and catalogue details - S-0012 / Tier B
PAL*M: Property Attestation for Large Generative Models ↗
P. Chantasantitam, A. I. Caulfield, V. Duddu, L. J. Gunn, N. Asokan · 2026 · arXiv
Supports: property attestation across training and inference on TDX + H100; formal model in Tamarin
Locator: abstract
Version and catalogue details - S-0018 / Tier B
A System Overview for Near-Term, Low-Trust AI Compute Verification ↗
N. Cankaya · 2026 · Machine Intelligence Research Institute
Supports: output cross-checks; inspector agents vulnerable to prompt injection
Locator: architecture; open problems
Version and catalogue details - S-1500 / Tier B
Proof-of-Guardrail in AI Agents and What (Not) to Trust from It ↗
X. Jin, M. Duan, Q. Lin, A. Chan, Z. Chen, J. Du, X. Ren · 2026 · arXiv
Supports: proof-of-guardrail protocol, prototype on AWS Nitro Enclaves with an external guardrail API, tamper tests, opt-in attestation, jailbreak risk
Locator: abstract; §3; §4.1; Table 1; Appendix A
Version and catalogue details - S-1503 / Tier B
Enabling Verifiably-Scoped Monitoring through Large Language Models and Trusted Compute ↗
B. Penchas, G. Zhao, R. Rinberg · 2026 · ICML 2026 Workshop on Technical AI Governance Research
Supports: monitoring scoped to a jointly signed plan and executed in an attested TEE
Locator: abstract
Version and catalogue details - S-1504 / Tier C
Auditor-in-a-Box: Tools for Third-Party Auditing ↗
R. Rinberg, B. Penchas · 2026 · LessWrong
Supports: Auditor-in-a-Box demo of plan-scoped monitoring; authors state user data and plan execution are not actually secure and the demo has not been stress-tested by a counterparty
Locator: whole post
Version and catalogue details
- Source review date
- 2026-09-25
- Drafted by (source map)
- ai
- Review handles (source map)
- codex-review