M-0012 / Cryptography & computation

Model identity attestation

Establishes that responses come from a specific, committed set of model weights, using enclave measurements or recomputation of sampled outputs.

R3 In productionSource reviewed 2026-09-25Provider-reported evidence

01 / The mechanism and its boundary

What the technique establishes

Model identity attestation shows users, auditors and regulators that a provider is serving the model weights it committed to. There are two routes. First, an attestation from a trusted execution environment shows that measured software enforced a hash commitment to the weights while the model ran. Second, a verifier holding the declared weights recomputes a sample of logged outputs, which can detect weights smuggled out in responses. Tinfoil reports running the enclave route commercially with its open-source Modelwrap tool. For unpublished weights, the commitment alone shows only that the same weights are served each time. Research prototypes bind evaluations to the same hash, so results apply to the served model. The enclave route inherits the weaknesses of TEEs, including published attacks that forge Intel TDX and AMD SEV-SNP attestations. Recomputation, tested on models of up to 30 billion parameters, needs trusted logging and randomness, and must tolerate numerical nondeterminism.

Threat model
Semi-trusted prover
Adversarial evaluation
Independent red-team
Hardware needed
Existing hardware features
Prover cooperation
Required
Confidentiality
Partial
Category
Cryptography & computation
Technical detail and cited results
  • Modelwrap build (Tinfoil's description). The tool downloads a pinned model revision, normalizes its directory structure and builds an EROFS image. It computes a dm-verity root hash with veritysetup. The Merkle tree gives a 32-byte commitment to, for example, 140 GB of weights S-0013.
  • Binding and enforcement. The root hash is passed on the kernel command line, which the enclave measurement covers. dm-verity then checks every block read against it and returns an error on any mismatch S-0013.
  • Reported costs. Tinfoil reports a hash-tree overhead of about 0.8% of image size, and build times from 5 s for a 549 MB model to 13 min 25 s for a 554 GB model. Cold-cache loading takes about 80% longer, and inference is unaffected once the weights are in GPU memory S-0013.
  • PAL*M inference attestation. PAL*M sets the Intel TDX REPORTDATA field to the concatenation of the operation, a challenge, and hashes of the inputs and outputs. It supports single-prompt and multi-turn inference attestations. Across three models of 3.8 to 8 billion parameters on an H100, PAL*M's added time was 3.8–11.4% of total run time for multi-turn sessions and 45.5–66.4% for single prompts, so a single attested prompt took 1.8 to 3 times as long as without PAL*M S-0012.
  • Fixed-seed sampling likelihood. Rinberg et al. score the likelihood that each token was honestly sampled from the claimed model under a known seed, with estimators for inverse-probability-transform and Gumbel-max samplers. On a 30B mixture-of-experts model under benign prompt traffic, the check limits exfiltratable information to under 0.5% at a false-positive rate under 0.01%, a slowdown of more than 200 times for the adversary S-0015. An independent study found that an adversary who controls the prompts roughly doubles leakage per token, reducing the slowdown to 60–118 times S-1507.

Claims and scope

A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.

Readiness for a stated use

R3 In production

Assessed use: showing users that a service runs the declared model weights

medium confidence · current · assessed 2026-09-25 · rubric 1.1

This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.

Tinfoil reports running the enclave route in a production service, but the independent attacks published so far left flaws open, including a critical flaw in the trusted execution environments that route relies on.

  • R1 met: designs and assumptions are published for both routes S-0013 S-0012 S-0015.
  • R2 met through Tinfoil's Modelwrap chain. Its code is open source S-1209, it runs on NVIDIA H100, H200 or B200 GPUs with AMD SEV-SNP or Intel TDX (provider-reported) S-1206, and it has been reported on models of up to 554 GB S-0013. Rinberg et al. publish code and results on models from 3B to 30B parameters S-0015.
  • R3 met on the provider's account: Tinfoil reports offering the feature in a production service S-1207 S-1208.
  • R4 not met, because the independent evaluations that exist left flaws open. TEE.fail used physical access to forge the Intel TDX attestation that the enclave route relies on. By pairing the forgeries with relayed H100 attestations, it made a vLLM proxy running outside TEE protection pass both the TDX and the NVIDIA confidential-computing checks S-1202. That flaw is critical and open. Kezins, at Delft University of Technology, attacked the exfiltration bound of Rinberg et al.'s recomputation check with chosen prompts. On a 30B mixture-of-experts model this raised leakage from 0.119 to 0.286 bits per token. The attack widens the covert channel and does not target the check that outputs match the declared model S-1507.

Evidence needed for the next level

  • Independent security evaluation of a deployed model-identity scheme that leaves no critical flaw open.

  • Resistance of the enclave variant to physical attackers (see TEE remote attestation for AI workloads).

  • Third-party verification for private models beyond consistency across requests.

  • Tooling for audit-time checking of transparency records.

Limitations, flaws, and blockers

These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.

critical / open / demonstrated attack

Underlying attestation can be forged or relayed

The enclave route inherits the platform-specific TEE attestation failures. Intel TDX forgery and H100 relay were demonstrated with physical access and host control S-1202. Battering RAM defeated AMD SEV-SNP attestation on DDR4 servers; RMPocalypse did so from malicious host software on platforms without AMD's fixes S-1210 S-1212 S-1213. These demonstrate failures of the trust roots, not of each model-commitment protocol.

S-1202S-1210S-1212S-1213S-1206S-0012
Response recorded by the source map

The TEE.fail authors report that physical interposer attacks are outside Intel's and AMD's threat models. AMD reports fixes for RMPocalypse S-1202 S-1213.

significant / mitigated / theoretical argument

Launch-state attestation does not by itself cover weights loaded later

Attestation measures launch state, and weights are read from disk after boot. A signature checked at load time does not stop a malicious hypervisor from altering the disk afterwards S-0013. Tinfoil reports mitigating this with dm-verity checks on every read S-0013. Unmeasured runtime configuration remains a general risk S-0014.

S-0013S-0014

significant / open / open question

For private models, a user can confirm consistency but not content

When weights are not published, users can check that the same root hash is served each time, but not what the model is S-0013. Pairing the hash with an attested evaluation, as in Attestable Audits, is one proposed remedy S-0009.

S-0013S-0009

significant / open / demonstrated attack

Recomputation depends on trusted logging and randomness, and its tolerance leaves a covert channel

The recomputation variant assumes that every input, output and seed is logged correctly, and that the attacker can neither predict nor manipulate which messages are sampled for verification. Legitimate nondeterminism concentrates at a few token positions, and slow leaks within the tolerated slack remain possible S-0015. An independent study showed that an adversary who controls the prompts roughly doubles the bits leaked per token, reducing the exfiltration slowdown from 146–254 times under benign prompts to 60–118 times. The attack targets the exfiltration bound, not the check that outputs match the declared model S-1507.

S-0015S-1507

What still blocks use or stronger assurance

  1. Attestation that resists physical attackers, for the enclave variant.

    Dependency: TEE remote attestation for AI workloads

    S-1202
  2. Numerical nondeterminism limits how tightly recomputation can pin down the model and sampling.

    Dependency: Deterministic and bit-exact inference

    S-0015
  3. The recomputation variant needs the verifier to hold the declared weights.

    S-0015

Connections in the research map

Complementary techniques

Alternative approaches

Concepts used

Organizations and developers

Implementations

The Consortium’s case files

Related editorial reviews use the Consortium’s own descriptive scores and review dates. Their scores are separate from the atlas readiness rubric.

RT-09 / AttestationTamperproof certificates, continuouslyRead case file ↗

Sources and provenance

  1. S-0013 / Tier C

    How Tinfoil Proves Exactly What Model Is Running ↗

    Tinfoil Team · 2026 · Tinfoil

    Supports: Modelwrap design, launch-state problem, signing comparison, private models, overheads (provider-reported)

    Locator: sections on the challenge, the three phases, performance, private models

    Version and catalogue details
  2. S-0012 / Tier B

    PAL*M: Property Attestation for Large Generative Models ↗

    P. Chantasantitam, A. I. Caulfield, V. Duddu, L. J. Gunn, N. Asokan · 2026 · arXiv

    Supports: inference attestation binding hashes to TDX report; overheads

    Locator: §4, Table 6

    Version and catalogue details
  3. S-0009 / Tier B

    Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments ↗

    C. Schnabl, D. Hugenroth, B. Marino, A. R. Beresford · 2025 · ICML 2025 Workshop on Technical AI Governance

    Supports: audit-to-inference model-hash binding; prototype evaluation of the audit step

    Locator: §3, Algorithms 1-3, §5

    Version and catalogue details
  4. S-0015 / Tier B

    Verifying LLM Inference to Detect Model Weight Exfiltration ↗

    R. Rinberg, A. Karvonen, A. Hoover, D. Reuter, K. Warr · 2025 · arXiv

    Supports: recomputation-based verification, assumptions, results, limitations, code

    Locator: Abstract; §1, §3.3-3.4, §4.2, §5-§7

    Version and catalogue details
  5. S-1507 / Tier B

    Adversarial Entropy Inflation Against Gumbel-Based Inference Verification ↗

    N. Kezins · 2026 · arXiv

    Supports: independent prompt-control attack that widens the exfiltration bound; 0.119 to 0.286 bits per token on the 30B MoE model

    Locator: Abstract; Table 1

    Version and catalogue details
  6. S-0014 / Tier C

    On TEEs for Privacy-Preserving Monitoring in AI Governance ↗

    Gloria Z · 2026 · MIRI Technical Governance Team

    Supports: measurement incompleteness; hashing-scheme warning

    Version and catalogue details
  7. S-3121 / Tier C

    What we learned about TEE security from auditing WhatsApp's Private Inference ↗

    Trail of Bits · 2026 · Trail of Bits blog

    Supports: Trail of Bits audit of WhatsApp Private Processing: environment variables loaded after the measurement (TOB-WAPI-13) and Meta's fix

    Version and catalogue details
  8. S-3124 / Tier B

    Meta WhatsApp Private Processing (security review) ↗

    Trail of Bits · 2025 · Trail of Bits publications library

    Supports: the review's finding that CVMs could be compromised through environment-variable injection

    Version and catalogue details
  9. S-1202 / Tier A

    TEE.fail: Breaking Trusted Execution Environments via DDR5 Memory Bus Interposition ↗

    J. Chuang, A. Seto, N. Berrios, S. van Schaik, C. Garman, D. Genkin · 2026 · 2026 IEEE Symposium on Security and Privacy (SP)

    Supports: Intel TDX attestation forgery and H100 attestation relay to a vLLM proxy outside TEE protection

    Locator: Abstract; §1.1, §8.3

    Version and catalogue details
  10. S-1210 / Tier A

    Battering RAM: Low-Cost Interposer Attacks on Confidential Computing via Dynamic Memory Aliasing ↗

    J. De Meulemeester, D. Oswald, I. Verbauwhede, J. Van Bulck · 2026 · 47th IEEE Symposium on Security and Privacy (S&P 2026)

    Supports: SEV-SNP attestation breach with a DDR4 interposer (Battering RAM)

    Locator: Abstract; site FAQ

    Version and catalogue details
  11. S-1212 / Tier A

    RMPocalypse: How a Catch-22 Breaks AMD SEV-SNP ↗

    B. Schlüter, S. Shinde · 2025 · 2025 ACM SIGSAC Conference on Computer and Communications Security (CCS '25)

    Supports: software-only SEV-SNP attestation forgery by a malicious hypervisor (RMPocalypse)

    Locator: Abstract; site

    Version and catalogue details
  12. S-1213 / Tier B

    SEV-SNP RMP Initialization Vulnerability (AMD-SB-3020) ↗

    AMD · 2025 · AMD product security bulletin

    Supports: AMD firmware fixes for RMPocalypse (vendor-reported)

    Locator: Mitigation tables

    Version and catalogue details
  13. S-1206 / Tier B

    A primer on secure enclaves ↗

    Tinfoil · 2026 · Tinfoil documentation

    Supports: Tinfoil hardware, trust model and documented limitations (provider-reported)

    Version and catalogue details
  14. S-1207 / Tier B

    Backend infrastructure ↗

    Tinfoil · 2026 · Tinfoil documentation

    Supports: measured boot chain, Sigstore publication, production model volumes (provider-reported)

    Version and catalogue details
  15. S-1208 / Tier B

    How verification works in Tinfoil ↗

    Tinfoil · 2026 · Tinfoil documentation

    Supports: production deployment; no supported audit-time tool (provider-reported)

    Locator: In-band vs. out-of-band verification

    Version and catalogue details
  16. S-1209 / Tier B

    modelwrap: Reproducible dm-verity read-only image of Huggingface models ↗

    Tinfoil · 2026 · GitHub

    Supports: open-source implementation, release v0.3.0

    Version and catalogue details
  17. S-1800 / Tier C

    Private Cloud Compute: A new frontier for AI privacy in the cloud ↗

    Apple Security Engineering and Architecture (SEAR) · 2024 · Apple Security Research blog

    Supports: Apple PCC: integrity protection of code and model assets; attestation against a public transparency log (provider-reported)

    Version and catalogue details
  18. S-1810 / Tier C

    Verde Verification System In Production ↗

    O. Ersoy · 2025 · Gensyn research blog

    Supports: Gensyn's statement that Verde shows outputs came from the declared model and data (provider-reported)

    Version and catalogue details
  19. S-1809 / Tier B

    Verde: Verification via Refereed Delegation for Machine Learning Programs ↗

    A. Arun, A. St. Arnaud, A. Titov, B. Wilcox, V. Kolobaric, M. Brinkmann, O. Ersoy, B. Fielding, J. Bonneau · 2025 · arXiv

    Supports: Verde's guarantee holds if at least one compute provider is honest

    Locator: Abstract

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review