01 / The mechanism and its boundary
What the technique establishes
Model identity attestation shows users, auditors and regulators that a provider is serving the model weights it committed to. There are two routes. First, an attestation from a trusted execution environment shows that measured software enforced a hash commitment to the weights while the model ran. Second, a verifier holding the declared weights recomputes a sample of logged outputs, which can detect weights smuggled out in responses. Tinfoil reports running the enclave route commercially with its open-source Modelwrap tool. For unpublished weights, the commitment alone shows only that the same weights are served each time. Research prototypes bind evaluations to the same hash, so results apply to the served model. The enclave route inherits the weaknesses of TEEs, including published attacks that forge Intel TDX and AMD SEV-SNP attestations. Recomputation, tested on models of up to 30 billion parameters, needs trusted logging and randomness, and must tolerate numerical nondeterminism.
- Threat model
- Semi-trusted prover
- Adversarial evaluation
- Independent red-team
- Hardware needed
- Existing hardware features
- Prover cooperation
- Required
- Confidentiality
- Partial
- Category
- Cryptography & computation
Technical detail and cited results
- Modelwrap build (Tinfoil's description). The tool downloads a pinned model revision, normalizes its directory structure and builds an EROFS image. It computes a dm-verity root hash with veritysetup. The Merkle tree gives a 32-byte commitment to, for example, 140 GB of weights S-0013.
- Binding and enforcement. The root hash is passed on the kernel command line, which the enclave measurement covers. dm-verity then checks every block read against it and returns an error on any mismatch S-0013.
- Reported costs. Tinfoil reports a hash-tree overhead of about 0.8% of image size, and build times from 5 s for a 549 MB model to 13 min 25 s for a 554 GB model. Cold-cache loading takes about 80% longer, and inference is unaffected once the weights are in GPU memory S-0013.
- PAL*M inference attestation. PAL*M sets the Intel TDX REPORTDATA field to the concatenation of the operation, a challenge, and hashes of the inputs and outputs. It supports single-prompt and multi-turn inference attestations. Across three models of 3.8 to 8 billion parameters on an H100, PAL*M's added time was 3.8–11.4% of total run time for multi-turn sessions and 45.5–66.4% for single prompts, so a single attested prompt took 1.8 to 3 times as long as without PAL*M S-0012.
- Fixed-seed sampling likelihood. Rinberg et al. score the likelihood that each token was honestly sampled from the claimed model under a known seed, with estimators for inverse-probability-transform and Gumbel-max samplers. On a 30B mixture-of-experts model under benign prompt traffic, the check limits exfiltratable information to under 0.5% at a false-positive rate under 0.01%, a slowdown of more than 200 times for the adversary S-0015. An independent study found that an adversary who controls the prompts roughly doubles leakage per token, reducing the slowdown to 60–118 times S-1507.
Claims and scope
A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.
The declared model is the one being served
Core purpose: responses come from the declared weights.
Declared safeguards were applied during inference
Links an attested evaluation to the model later served (S-0009).
Model weights have not left the facility
The recomputation variant limits steganographic weight exfiltration through outputs (S-0015).
Readiness for a stated use
Assessed use: showing users that a service runs the declared model weights
medium confidence · current · assessed 2026-09-25 · rubric 1.1
This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.
Tinfoil reports running the enclave route in a production service, but the independent attacks published so far left flaws open, including a critical flaw in the trusted execution environments that route relies on.
- R1 met: designs and assumptions are published for both routes S-0013 S-0012 S-0015.
- R2 met through Tinfoil's Modelwrap chain. Its code is open source S-1209, it runs on NVIDIA H100, H200 or B200 GPUs with AMD SEV-SNP or Intel TDX (provider-reported) S-1206, and it has been reported on models of up to 554 GB S-0013. Rinberg et al. publish code and results on models from 3B to 30B parameters S-0015.
- R3 met on the provider's account: Tinfoil reports offering the feature in a production service S-1207 S-1208.
- R4 not met, because the independent evaluations that exist left flaws open. TEE.fail used physical access to forge the Intel TDX attestation that the enclave route relies on. By pairing the forgeries with relayed H100 attestations, it made a vLLM proxy running outside TEE protection pass both the TDX and the NVIDIA confidential-computing checks S-1202. That flaw is critical and open. Kezins, at Delft University of Technology, attacked the exfiltration bound of Rinberg et al.'s recomputation check with chosen prompts. On a 30B mixture-of-experts model this raised leakage from 0.119 to 0.286 bits per token. The attack widens the covert channel and does not target the check that outputs match the declared model S-1507.
Evidence needed for the next level
Independent security evaluation of a deployed model-identity scheme that leaves no critical flaw open.
Resistance of the enclave variant to physical attackers (see TEE remote attestation for AI workloads).
Third-party verification for private models beyond consistency across requests.
Tooling for audit-time checking of transparency records.
Limitations, flaws, and blockers
These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.
critical / open / demonstrated attack
Underlying attestation can be forged or relayed
The enclave route inherits the platform-specific TEE attestation failures. Intel TDX forgery and H100 relay were demonstrated with physical access and host control S-1202. Battering RAM defeated AMD SEV-SNP attestation on DDR4 servers; RMPocalypse did so from malicious host software on platforms without AMD's fixes S-1210 S-1212 S-1213. These demonstrate failures of the trust roots, not of each model-commitment protocol.
significant / mitigated / theoretical argument
Launch-state attestation does not by itself cover weights loaded later
Attestation measures launch state, and weights are read from disk after boot. A signature checked at load time does not stop a malicious hypervisor from altering the disk afterwards S-0013. Tinfoil reports mitigating this with dm-verity checks on every read S-0013. Unmeasured runtime configuration remains a general risk S-0014.
significant / open / open question
For private models, a user can confirm consistency but not content
When weights are not published, users can check that the same root hash is served each time, but not what the model is S-0013. Pairing the hash with an attested evaluation, as in Attestable Audits, is one proposed remedy S-0009.
significant / open / demonstrated attack
Recomputation depends on trusted logging and randomness, and its tolerance leaves a covert channel
The recomputation variant assumes that every input, output and seed is logged correctly, and that the attacker can neither predict nor manipulate which messages are sampled for verification. Legitimate nondeterminism concentrates at a few token positions, and slow leaks within the tolerated slack remain possible S-0015. An independent study showed that an adversary who controls the prompts roughly doubles the bits leaked per token, reducing the exfiltration slowdown from 146–254 times under benign prompts to 60–118 times. The attack targets the exfiltration bound, not the check that outputs match the declared model S-1507.
What still blocks use or stronger assurance
Attestation that resists physical attackers, for the enclave variant.
Dependency: TEE remote attestation for AI workloads
S-1202Numerical nondeterminism limits how tightly recomputation can pin down the model and sampling.
Dependency: Deterministic and bit-exact inference
S-0015- S-0015
The recomputation variant needs the verifier to hold the declared weights.
Connections in the research map
Complementary techniques
Alternative approaches
Concepts used
Organizations and developers
Implementations
The Consortium’s case files
Related editorial reviews use the Consortium’s own descriptive scores and review dates. Their scores are separate from the atlas readiness rubric.
RT-09 / AttestationTamperproof certificates, continuouslyRead case file ↗Sources and provenance
- S-0013 / Tier C
How Tinfoil Proves Exactly What Model Is Running ↗
Tinfoil Team · 2026 · Tinfoil
Supports: Modelwrap design, launch-state problem, signing comparison, private models, overheads (provider-reported)
Locator: sections on the challenge, the three phases, performance, private models
Version and catalogue details - S-0012 / Tier B
PAL*M: Property Attestation for Large Generative Models ↗
P. Chantasantitam, A. I. Caulfield, V. Duddu, L. J. Gunn, N. Asokan · 2026 · arXiv
Supports: inference attestation binding hashes to TDX report; overheads
Locator: §4, Table 6
Version and catalogue details - S-0009 / Tier B
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments ↗
C. Schnabl, D. Hugenroth, B. Marino, A. R. Beresford · 2025 · ICML 2025 Workshop on Technical AI Governance
Supports: audit-to-inference model-hash binding; prototype evaluation of the audit step
Locator: §3, Algorithms 1-3, §5
Version and catalogue details - S-0015 / Tier B
Verifying LLM Inference to Detect Model Weight Exfiltration ↗
R. Rinberg, A. Karvonen, A. Hoover, D. Reuter, K. Warr · 2025 · arXiv
Supports: recomputation-based verification, assumptions, results, limitations, code
Locator: Abstract; §1, §3.3-3.4, §4.2, §5-§7
Version and catalogue details - S-1507 / Tier B
Adversarial Entropy Inflation Against Gumbel-Based Inference Verification ↗
N. Kezins · 2026 · arXiv
Supports: independent prompt-control attack that widens the exfiltration bound; 0.119 to 0.286 bits per token on the 30B MoE model
Locator: Abstract; Table 1
Version and catalogue details - S-0014 / Tier C
On TEEs for Privacy-Preserving Monitoring in AI Governance ↗
Gloria Z · 2026 · MIRI Technical Governance Team
Supports: measurement incompleteness; hashing-scheme warning
Version and catalogue details - S-3121 / Tier C
What we learned about TEE security from auditing WhatsApp's Private Inference ↗
Trail of Bits · 2026 · Trail of Bits blog
Supports: Trail of Bits audit of WhatsApp Private Processing: environment variables loaded after the measurement (TOB-WAPI-13) and Meta's fix
Version and catalogue details - S-3124 / Tier B
Meta WhatsApp Private Processing (security review) ↗
Trail of Bits · 2025 · Trail of Bits publications library
Supports: the review's finding that CVMs could be compromised through environment-variable injection
Version and catalogue details - S-1202 / Tier A
TEE.fail: Breaking Trusted Execution Environments via DDR5 Memory Bus Interposition ↗
J. Chuang, A. Seto, N. Berrios, S. van Schaik, C. Garman, D. Genkin · 2026 · 2026 IEEE Symposium on Security and Privacy (SP)
Supports: Intel TDX attestation forgery and H100 attestation relay to a vLLM proxy outside TEE protection
Locator: Abstract; §1.1, §8.3
Version and catalogue details - S-1210 / Tier A
Battering RAM: Low-Cost Interposer Attacks on Confidential Computing via Dynamic Memory Aliasing ↗
J. De Meulemeester, D. Oswald, I. Verbauwhede, J. Van Bulck · 2026 · 47th IEEE Symposium on Security and Privacy (S&P 2026)
Supports: SEV-SNP attestation breach with a DDR4 interposer (Battering RAM)
Locator: Abstract; site FAQ
Version and catalogue details - S-1212 / Tier A
RMPocalypse: How a Catch-22 Breaks AMD SEV-SNP ↗
B. Schlüter, S. Shinde · 2025 · 2025 ACM SIGSAC Conference on Computer and Communications Security (CCS '25)
Supports: software-only SEV-SNP attestation forgery by a malicious hypervisor (RMPocalypse)
Locator: Abstract; site
Version and catalogue details - S-1213 / Tier B
SEV-SNP RMP Initialization Vulnerability (AMD-SB-3020) ↗
AMD · 2025 · AMD product security bulletin
Supports: AMD firmware fixes for RMPocalypse (vendor-reported)
Locator: Mitigation tables
Version and catalogue details - S-1206 / Tier B
A primer on secure enclaves ↗
Tinfoil · 2026 · Tinfoil documentation
Supports: Tinfoil hardware, trust model and documented limitations (provider-reported)
Version and catalogue details - S-1207 / Tier B
Backend infrastructure ↗
Tinfoil · 2026 · Tinfoil documentation
Supports: measured boot chain, Sigstore publication, production model volumes (provider-reported)
Version and catalogue details - S-1208 / Tier B
How verification works in Tinfoil ↗
Tinfoil · 2026 · Tinfoil documentation
Supports: production deployment; no supported audit-time tool (provider-reported)
Locator: In-band vs. out-of-band verification
Version and catalogue details - S-1209 / Tier B
modelwrap: Reproducible dm-verity read-only image of Huggingface models ↗
Tinfoil · 2026 · GitHub
Supports: open-source implementation, release v0.3.0
Version and catalogue details - S-1800 / Tier C
Private Cloud Compute: A new frontier for AI privacy in the cloud ↗
Apple Security Engineering and Architecture (SEAR) · 2024 · Apple Security Research blog
Supports: Apple PCC: integrity protection of code and model assets; attestation against a public transparency log (provider-reported)
Version and catalogue details - S-1810 / Tier C
Verde Verification System In Production ↗
O. Ersoy · 2025 · Gensyn research blog
Supports: Gensyn's statement that Verde shows outputs came from the declared model and data (provider-reported)
Version and catalogue details - S-1809 / Tier B
Verde: Verification via Refereed Delegation for Machine Learning Programs ↗
A. Arun, A. St. Arnaud, A. Titov, B. Wilcox, V. Kolobaric, M. Brinkmann, O. Ersoy, B. Fielding, J. Bonneau · 2025 · arXiv
Supports: Verde's guarantee holds if at least one compute provider is honest
Locator: Abstract
Version and catalogue details
- Source review date
- 2026-09-25
- Drafted by (source map)
- ai
- Review handles (source map)
- codex-review