C-0005

The declared model is the one being served

Outputs delivered to users or auditors come from the specific model, weights and configuration the provider declared, not from a substitute.

Source reviewed 2026-09-25

01 / The mechanism and its boundary

What is being described

The claim is that outputs delivered to users or auditors come from the specific model, weights and configuration the provider declared. Evaluations, audits and agreements often apply to one specific model. If a provider could evaluate one model and serve another, such as a cheaper, quantized or modified version, those checks would say little about what users receive. It is a positive claim that can be tested directly, but three problems make it hard. Numerical nondeterminism means honest recomputation does not match exactly. The verifier usually cannot see the weights, which are commercially or strategically sensitive. And the evidence must come from the actual serving system rather than a separate test instance. Approaches include statistical or exact recomputation of sampled outputs, hardware attestation of the loaded weights, and zero-knowledge proofs. They trade off cost, trust in hardware vendors and confidentiality.

State of verification

Editorial synthesis from the AI Verification Tech Map.

The served model can be checked directly on its outputs, and several mechanisms for it are demonstrated (R2) or in production (R3). None is deployment-ready (R4): independent security evaluations are either missing or, for trusted execution environments, found a critical flaw.

For example, users of a hosted model may want to know that the model answering them is the one an auditor evaluated.

  • Model identity attestation (R3) extends TEE remote attestation (R3) from the software an enclave booted to the weights it loads. Confidential multi-party verification (R2) runs the audit in an enclave and binds its result to the model's hash S-0009.
  • Attestable Audits (R2) joins the two steps. At inference, the enclave checks the served model's hash against the audited one and returns each response with an attestation that links model, prompt, output and audit result S-0009. Its reported evaluation covers only the audit step, on CPU-only enclaves with a 4-bit, 8-billion-parameter model S-0009. Tinfoil's Modelwrap (R3) covers serving alone, in what Tinfoil reports is a production service S-1208.
  • The evidence is narrower than what users want. Attested enclaves report that the weights that answered have the same hash as the weights that were audited. For a public model, Modelwrap lets anyone rebuild the hash from the published weights. For a private model, users see only the hash S-0013.
  • The chain trusts the TEE vendors' attestation keys. With physical access, researchers forged Intel TDX attestations and paired them with relayed H100 attestations, so that a system outside TEE protection passed both checks S-1202. Sampled recomputation, as in DiFR (R2), avoids hardware trust but needs a trusted copy of the weights and evidence tied to the production serving path S-0018. Zero-knowledge proofs of inference (R2) keep the weights private without trusting hardware, but proving remains expensive S-0018.

Connections in the research map

Concepts used

Techniques addressing this claim

Sources and provenance

  1. S-0002 / Tier B

    Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment ↗

    M. Baker, G. Kulp, O. Marks, M. Brundage, L. Heim · 2025 · RAND Corporation

    Supports: Subgoal 1.A (declared uses declared accurately, including inference) and 1.B (required properties; deployed models evaluated at intervals)

    Locator: §3.2

    Version and catalogue details
  2. S-0016 / Tier B

    DiFR: Inference Verification Despite Nondeterminism ↗

    A. Karvonen, D. Reuter, R. Rinberg, L. Marks, A. Garriga-Alonso, K. Warr · 2025 · ICML 2026 Workshop on Technical AI Governance Research

    Supports: need to verify inference; nondeterminism; Token-DiFR detects 4-bit quantization with AUC > 0.999 within 300 tokens

    Locator: abstract

    Version and catalogue details
  3. S-1009 / Tier C

    Defeating Nondeterminism in LLM Inference ↗

    H. He, Thinking Machines Lab · 2025 · Thinking Machines Lab: Connectionism

    Supports: batch-size dependence as a cause of inference nondeterminism

    Locator: batch invariance section

    Version and catalogue details
  4. S-0020 / Tier B

    Bit-Exact AI Inference Verification Without Performance Tradeoffs ↗

    N. Cankaya · 2026 · ICML 2026 Workshop on Technical AI Governance Research

    Supports: bit-exact reproduction across GPU variants given recomputation data

    Locator: abstract

    Version and catalogue details
  5. S-0004 / Tier B

    Verification for International AI Governance ↗

    B. Harack, R. F. Trager, A. Reuel, D. Manheim, M. Brundage, O. Aarne, A. Scher, Y. Pan, J. Xiao, K. Loke, S. N. Adan, G. Bas, N. A. Caputo, J. C. Morse, J. Ahuja, I. Duan, J. Egan, B. Bucknall, B. Rosen, R. Araujo, V. Boulanin, R. Lall, F. Barez, S. Alvira, C. Katzke, A. Atamli, A. Awad · 2025 · Oxford Martin AI Governance Initiative

    Supports: model fingerprint attestation; device-model mating with an encrypted model

    Locator: Appendix K (p. 157); Appendix L.4 (p. 159)

    Version and catalogue details
  6. S-0018 / Tier B

    A System Overview for Near-Term, Low-Trust AI Compute Verification ↗

    N. Cankaya · 2026 · Machine Intelligence Research Institute

    Supports: whitelisted models for blacklisted uses; attributing forward passes to hardware and time; committed weights in auditing environments; ZKP cost

    Locator: verification goals; architecture; open problems

    Version and catalogue details
  7. S-0023 / Tier A

    zkLLM: Zero Knowledge Proofs for Large Language Models ↗

    H. Sun, J. Li, H. Zhang · 2024 · 2024 ACM SIGSAC Conference on Computer and Communications Security (CCS 2024)

    Supports: zkLLM proves one 2,048-token forward pass of a 13B-parameter model in under 15 minutes with proofs under 200 kB, keeping parameters private

    Locator: abstract; §8 Table 1

    Version and catalogue details
  8. S-0012 / Tier B

    PAL*M: Property Attestation for Large Generative Models ↗

    P. Chantasantitam, A. I. Caulfield, V. Duddu, L. J. Gunn, N. Asokan · 2026 · arXiv

    Supports: property attestation on Intel TDX + NVIDIA H100 with under 11% overhead for common operations

    Locator: abstract

    Version and catalogue details
  9. S-0014 / Tier C

    On TEEs for Privacy-Preserving Monitoring in AI Governance ↗

    Gloria Z · 2026 · MIRI Technical Governance Team

    Supports: attestation-key holder can produce valid reports; side-channel and physical attacks; measurement coverage

    Locator: Limitations

    Version and catalogue details
  10. S-0009 / Tier B

    Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments ↗

    C. Schnabl, D. Hugenroth, B. Marino, A. R. Beresford · 2025 · ICML 2025 Workshop on Technical AI Governance

    Supports: attestation linking model, audit result, prompt and response in a TEE; reported evaluation covers the audit step on CPU-only enclaves with a 4-bit 8B model

    Locator: abstract; inference protocol; §5

    Version and catalogue details
  11. S-0013 / Tier C

    How Tinfoil Proves Exactly What Model Is Running ↗

    Tinfoil Team · 2026 · Tinfoil

    Supports: public models' hashes can be rebuilt; private models expose only the hash (provider-reported)

    Version and catalogue details
  12. S-1208 / Tier B

    How verification works in Tinfoil ↗

    Tinfoil · 2026 · Tinfoil documentation

    Supports: Modelwrap chain in Tinfoil's production service (provider-reported)

    Locator: In-band vs. out-of-band verification

    Version and catalogue details
  13. S-1202 / Tier A

    TEE.fail: Breaking Trusted Execution Environments via DDR5 Memory Bus Interposition ↗

    J. Chuang, A. Seto, N. Berrios, S. van Schaik, C. Garman, D. Genkin · 2026 · 2026 IEEE Symposium on Security and Privacy (SP)

    Supports: physical memory-bus interposition extracts a per-CPU Intel attestation key and forges TDX attestations

    Locator: abstract; §1.1

    Version and catalogue details
  14. S-1210 / Tier A

    Battering RAM: Low-Cost Interposer Attacks on Confidential Computing via Dynamic Memory Aliasing ↗

    J. De Meulemeester, D. Oswald, I. Verbauwhede, J. Van Bulck · 2026 · 47th IEEE Symposium on Security and Privacy (S&P 2026)

    Supports: Battering RAM forges SEV-SNP attestation with a DDR4 interposer

    Version and catalogue details
  15. S-1212 / Tier A

    RMPocalypse: How a Catch-22 Breaks AMD SEV-SNP ↗

    B. Schlüter, S. Shinde · 2025 · 2025 ACM SIGSAC Conference on Computer and Communications Security (CCS '25)

    Supports: RMPocalypse forges SEV-SNP attestation from a malicious hypervisor

    Version and catalogue details
  16. S-3540 / Tier A

    Model Equality Testing: Which Model Is This API Serving? ↗

    I. Gao, P. Liang, C. Guestrin · 2025 · International Conference on Learning Representations (ICLR 2025)

    Supports: model equality testing: median 77.4% power with about 10 samples per prompt; 11 of 31 Llama API endpoints in summer 2024 served distributions different from the reference weights

    Locator: abstract

    Version and catalogue details
  17. S-3541 / Tier B

    Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs ↗

    W. Cai, T. Shi, X. Zhao, D. Song · 2025 · arXiv

    Supports: output tests query-intensive and fail against subtle substitutions; log-probability tests defeated by inference nondeterminism

    Locator: abstract

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review