C-0006

Declared safeguards were applied during inference

Specified safety measures, such as input filters, output checks or monitoring, actually ran on the requests a deployed model served.

Source reviewed 2026-09-25

01 / The mechanism and its boundary

What is being described

The claim is that specified safety measures, such as input filters, output checks or monitoring for misuse, actually ran on the requests a deployed model served. Rules for deploying AI systems often require such safeguards. A developer's statement that safeguards exist does not show that they ran on every request, or that the version checked is the one in production. Verifying application matters to regulators and to agreements that allow deployment on condition of mitigations. It is a positive claim, but it inherits the difficulties of verifying the served model and adds more. Safeguards are often separate software components whose configuration can change, and the verifier usually cannot see users' prompts or the provider's systems. Proposals use trusted execution environments to attest the serving stack, audits run inside such environments, and inspections. Whether a safeguard is effective is a separate question from whether it was applied.

State of verification

Editorial synthesis from the AI Verification Tech Map.

Safeguard application has been shown only in small research prototypes, which can show that a declared safeguard ran, not that it works.

Safeguard attestation (R2) runs the safeguard inside a trusted execution environment, whose hardware signs a measurement of its code with a commitment to each input and response S-1500. It builds on TEE remote attestation (R3) and on knowing which model is served (The declared model is the one being served). The proof-of-guardrail prototype demonstrates it on CPU enclaves in AWS and calls its guardrail model through an external API S-1500. In the authors' tests it detected modified guardrail code, attestations and responses S-1500. Confidential multi-party verification (R2) can limit monitoring to a plan both parties sign S-1503, and Auditor-in-a-Box demonstrates such plan-scoped monitoring, though its authors state that the demo's user data and plan execution are not actually secure S-1504.

Every component that influences inference must be covered by the launch measurement S-0014, and the prototype attests only the responses for which it offers attestation, so coverage of all traffic is not shown S-1500. Where hardware trust is unavailable, the sources fall back on inspections, audits and personnel-based layers S-0062 S-0003 S-0002.

Connections in the research map

Concepts used

Techniques addressing this claim

Sources and provenance

  1. S-0002 / Tier B

    Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment ↗

    M. Baker, G. Kulp, O. Marks, M. Brundage, L. Heim · 2025 · RAND Corporation

    Supports: deployment mitigations specified by inputs and outputs (filters, oversight checks); difficulty of choosing technical rules; whistleblower and interview layers

    Locator: Table 3; §4

    Version and catalogue details
  2. S-0062 / Tier B

    Verification methods for international AI agreements ↗

    A. R. Wasil, T. Reed, J. W. Miller, P. Barnett · 2024 · arXiv

    Supports: AI developer inspections to check authorised code and implementation of evaluations and safeguards; software can be quickly modified or hidden

    Locator: Access-dependent methods; Table 1

    Version and catalogue details
  3. S-0003 / Tier B

    Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies ↗

    M. Brundage, N. Dreksler, A. Homewood, S. McGregor, P. Paskov, C. Stosz, G. Sastry, A. F. Cooper, G. Balston, S. Adler, S. Casper, M. Anderljung, G. Werner, S. Mindermann, V. Mavroudis, B. Bucknall, C. Stix, J. Freund, L. Pacchiardi, J. Hernandez-Orallo, M. Pistillo, M. Chen, C. Painter, D. W. Ball, C. O'Keefe, G. Weil, B. Harack, G. Finley, R. Hassan, S. Emmons, C. Foster, A. Reuel, B. Treece, Y. Bengio, D. Reti, R. Bommasani, C. Trout, A. S. Shamsabadi, R. Dattani, A. Weller, R. Trager, J. Sevilla, L. Wagner, L. Soder, K. Ramakrishnan, H. Papadatos, M. Murray, R. Tovcimak · 2026 · arXiv

    Supports: third-party verification of developers' safety and security claims with deep, secure access; AI Assurance Levels

    Locator: abstract

    Version and catalogue details
  4. S-0014 / Tier C

    On TEEs for Privacy-Preserving Monitoring in AI Governance ↗

    Gloria Z · 2026 · MIRI Technical Governance Team

    Supports: TEEs for policy-adherent inference; launch measurement must cover all components; attestation-key holder can forge reports

    Locator: main argument; Limitations

    Version and catalogue details
  5. S-0009 / Tier B

    Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments ↗

    C. Schnabl, D. Hugenroth, B. Marino, A. R. Beresford · 2025 · ICML 2025 Workshop on Technical AI Governance

    Supports: TEE-based verifiable benchmarks keeping model IP and datasets confidential; prototype with Llama-3.1

    Locator: abstract

    Version and catalogue details
  6. S-0012 / Tier B

    PAL*M: Property Attestation for Large Generative Models ↗

    P. Chantasantitam, A. I. Caulfield, V. Duddu, L. J. Gunn, N. Asokan · 2026 · arXiv

    Supports: property attestation across training and inference on TDX + H100; formal model in Tamarin

    Locator: abstract

    Version and catalogue details
  7. S-0018 / Tier B

    A System Overview for Near-Term, Low-Trust AI Compute Verification ↗

    N. Cankaya · 2026 · Machine Intelligence Research Institute

    Supports: output cross-checks; inspector agents vulnerable to prompt injection

    Locator: architecture; open problems

    Version and catalogue details
  8. S-1500 / Tier B

    Proof-of-Guardrail in AI Agents and What (Not) to Trust from It ↗

    X. Jin, M. Duan, Q. Lin, A. Chan, Z. Chen, J. Du, X. Ren · 2026 · arXiv

    Supports: proof-of-guardrail protocol, prototype on AWS Nitro Enclaves with an external guardrail API, tamper tests, opt-in attestation, jailbreak risk

    Locator: abstract; §3; §4.1; Table 1; Appendix A

    Version and catalogue details
  9. S-1503 / Tier B

    Enabling Verifiably-Scoped Monitoring through Large Language Models and Trusted Compute ↗

    B. Penchas, G. Zhao, R. Rinberg · 2026 · ICML 2026 Workshop on Technical AI Governance Research

    Supports: monitoring scoped to a jointly signed plan and executed in an attested TEE

    Locator: abstract

    Version and catalogue details
  10. S-1504 / Tier C

    Auditor-in-a-Box: Tools for Third-Party Auditing ↗

    R. Rinberg, B. Penchas · 2026 · LessWrong

    Supports: Auditor-in-a-Box demo of plan-scoped monitoring; authors state user data and plan execution are not actually secure and the demo has not been stress-tested by a counterparty

    Locator: whole post

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review