01 / The mechanism and its boundary
What the technique establishes
Confidential multi-party verification runs an agreed check over assets that their owners will not show anyone, such as a developer's model weights, an auditor's test data or users' logs, so that no participant can see the others' inputs. The check runs either inside hardware-isolated enclaves that sign what code ran or with zero-knowledge proofs. Only the result is released, with evidence of how it was produced. In a 2026 cloud GPU pilot, an outside evaluator tested a proprietary Gemini model on private prompts, neither side seeing the other's inputs. Research prototypes compose multi-step audit workflows, limit usage monitoring to a jointly signed plan, and prove audits of small models with zero-knowledge proofs. No independent red-team of these systems has been published. The main obstacles are scaling to frontier models, and trust in hardware vendors and the cloud host. Even a one-bit verdict can leak information about the private inputs.
- Threat model
- Semi-trusted prover
- Adversarial evaluation
- Published analysis
- Hardware needed
- Existing hardware features
- Prover cooperation
- Required
- Confidentiality
- Preserving
- Category
- Cryptography & computation
Technical detail and cited results
- Cove object model. An artifact is a unit of private data with exactly one owner, encrypted locally before upload. A workflow is a directed acyclic graph with a canonical serialization, and its hash names the bundle. Compilation to per-node manifests is deterministic, and the manifest hash is the node's identity. Owners record allow rules that bind each artifact to approved manifest hashes S-1505.
- Cove run and verification. At run time a node fetches and verifies upstream certificates, evaluates preconditions, obtains decryption keys only after attestation, runs the workload and emits a certificate whose TEE report data commits to the certificate-body hash. Report data is two SHA-256 hashes, of a fixed label and of a canonical payload. Verifiers start from a terminal certificate, the workflow bytes and the TEE vendor's roots, and check each certificate recursively S-1505.
- ZkAudit. The provider publishes commitments to its dataset and weights, with a zero-knowledge proof that the committed weights result from training. It then answers each audit request by computing a function F privately and releasing F's output with a second proof S-0022. A copyright or demographic audit of a MobileNet v2 model on Flowers-102 cost $108 in total (about 10 cents per image), and a counterfactual audit of a recommender on MovieLens cost $8,456 S-0022.
- Minimal Information Disclosure. It chooses the evidence mechanism to minimize the conditional mutual information between protected properties and the released evidence, subject to a floor on information about the authorized target. One variant proves a linear projection with a Groth16 zk-SNARK S-1506.
Claims and scope
A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.
The declared evaluation was run (upstream draft)
Attestable Audits binds the model, audit code and data, and result. PySyft's pilot ran an agreed private evaluation with participant-accepted opaque-code and cloud trust assumptions (S-0009, S-3320).
The declared model is the one being served
Binds audit or capability-evaluation results to the model that is served, without revealing weights (S-0009, S-0011).
Declared safeguards were applied during inference
Plan-scoped monitoring runs an agreed classifier over private usage records (S-1503).
A training run stayed within declared limits
Zero-knowledge audits can prove properties of committed training data and weights (S-0022).
Readiness for a stated use
Assessed use: audits or evaluations of a private model that reveal neither party's inputs
medium confidence · current · assessed 2026-09-28 · rubric 1.1
This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.
A 2026 pilot evaluated a proprietary model on a generally available cloud enclave service, but the evaluation workflows are pilots or research prototypes, and no party is documented as relying on their results for a verification decision.
- R1 met: designs that state what is verified and what is trusted are published for TEE workflows S-0011 S-1505 and zero-knowledge audits S-0022.
- R2 met. In the double-blind pilot, AVERI evaluated Gemini 2.5 Flash Lite against private benchmark prompts in a Google Cloud enclave with an NVIDIA H100, with the prompts kept from Google DeepMind and the weights from the evaluators (participant-reported) S-3320. Cove has an open-source reference implementation on Intel TDX via Phala Cloud's dstack, demonstrated end to end on a benchmark workflow S-0011 S-1505. ZkAudit reports peer-reviewed end-to-end audits of MobileNet image classifiers and a recommender model S-0022. Attestable Audits ran safety benchmarks on an 8-billion-parameter model in AWS Nitro Enclaves S-0009.
- R3 not met for this use. Google reports that Confidential Space, which releases each data owner's data only to an attested workload that meets the owner's conditions, is generally available, on H100 GPUs since April 2026 S-3321 S-3322. That release step is TEE attestation, assessed on its own record. The evaluation workflow that ran on it, OpenMined's PySyft, is documented only as a pilot S-3320. Pour Demain, an outside auditor, ran interpretability evaluations of a 744-billion-parameter open-weights model on Tinfoil's production confidential-computing platform, with its governance layer shown for single sessions only S-3361. No party is documented as relying on such a result for a verification decision.
- R4 not met: no independent audit or red-team of Confidential Space for this use, Cove, Attestable Audits or ZkAudit has been published. An independent analysis of one monitoring tool's published evaluation found that its one-bit verdicts leak information S-1506. Attackers with physical access to the memory bus have forged the Intel TDX attestation the pilot relied on S-1202 S-3126.
Confidence is medium. The R2 evidence includes peer-reviewed results S-0022 and a pilot on a proprietary model S-3320. Whether a general-purpose enclave service can carry this use to R3 is a judgment call.
Evidence needed for the next level
A production-grade, available confidential evaluation or audit workflow, or reliance by a party other than its developer on such a result for a verification decision.
An independent public evaluation (audit, red-team or peer-reviewed security analysis) of a confidential multi-party verification system, such as Confidential Space with PySyft or Cove, that leaves no critical flaw open.
Confidential evaluation on multi-node enclave clusters at the scale of the largest frontier models, or zero-knowledge audits at that scale.
Limitations, flaws, and blockers
These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.
significant / open / demonstrated attack
Released verdicts can leak information about private inputs
Even a one-bit result can reveal more than intended. Abdelghafar and Kulp used the published evaluation of Auditor-in-a-Box, whose output filter is meant to disclose at most one bit. Given only the valid or invalid decision on a new request, a simple predictor identified which of two request sets it came from (borderline or proxy requests versus ordinary research requests) with 72.4% balanced accuracy, against 50% without the decision. In a second experiment, which distinguished reinforcement-learning workloads from others, several candidate one-bit reports revealed exact-workload information.
significant / open / demonstrated attack
Memory-bus interposition extracts attestation keys and forges attestations
The platform-specific attacks and fixes are collected in the TEE findings. Intel TDX attestations were forged with physical access to DDR5 systems and combined with H100 relay S-1202 S-3126. AMD SEV-SNP attestation was defeated on DDR4 by Battering RAM and from malicious host software by RMPocalypse before AMD's fixes S-1210 S-1212 S-1213.
significant / open / theoretical argument
Guarantees depend on the host software stack and on review
Cove's developers state that compromise of the Docker daemon, host kernel or TEE stack breaks all guarantees. They also state that Docker policy alone cannot prove that guest code cannot generate a quote if the platform exposes quote instructions globally, and that compiled workflow bundles are hashed and reviewable but not signed by a publisher key.
significant / open / theoretical argument
Completeness is not established
A confidential workflow proves facts about the records and models submitted to it. A governance analysis notes that an auditor also needs assurance that all activity is accounted for, since a host could start a second confidential virtual machine that uses a different model or bypasses monitoring.
minor / open / theoretical argument
Zero-knowledge audits reveal model architecture
ZkAudit keeps weights and data secret but reveals the model architecture, and it does not protect against data poisoning.
What still blocks use or stronger assurance
Frontier inference typically needs the resources of several GPUs, and published confidential evaluations ran on a single server, so the double-blind pilot's authors name many-node confidential H100 or B200 clusters as the next milestone.
Dependency: TEE remote attestation for AI workloads
S-0014S-3361S-3320Zero-knowledge audits have been shown on image classifiers and a recommender model, not language models at frontier scale, and a counterfactual audit of the recommender cost $8,456.
Dependency: Zero-knowledge proofs of inference
S-0022Trust rests on a small number of hardware vendors, and a per-CPU Intel attestation key has been extracted by physical attack.
Dependency: TEE remote attestation for AI workloads
S-0014S-1202- S-1504
Parties must negotiate the plan or workflow and handle false positives and appeals, which the Auditor-in-a-Box authors list as open problems.
- S-1506
Released evidence must be designed to limit collateral leakage, which requires declaring protected properties in advance and calibrating on labelled executions.
Connections in the research map
Depends on
- TEE remote attestation for AI workloads
TEE-based designs rely on measured launch and remote attestation.
Complementary techniques
Concepts used
Organizations and developers
Implementations
The Consortium’s case files
Related editorial reviews use the Consortium’s own descriptive scores and review dates. Their scores are separate from the atlas readiness rubric.
WV-06 / WorkloadInspect everything, reveal nothingRead case file ↗Sources and provenance
- S-0011 / Tier B
Cove: Compositional Multi-Party Confidential Workflows for Verifiable AI Governance ↗
S. Ding, E. Lee, R. Cheng, D. Kang · 2026 · ICML 2026 Workshop on Technical AI Governance Research
Supports: problem statement; framework; three applications; open-source implementation on Intel TDX via dstack
Locator: abstract (read via the ICML 2026 virtual poster page; the OpenReview PDF was not reachable)
Version and catalogue details - S-1505 / Tier B
Cove: Compositional Multi-Party Confidential Workflows for Verifiable AI Governance (reference implementation) ↗
covehub · 2026 · GitHub
Supports: object model; lifecycle; certificates; trust boundary; residual risks
Locator: README; docs/internal/architecture.md; docs/internal/security_model.md
Version and catalogue details - S-0022 / Tier A
Trustless Audits without Revealing Data or Models ↗
S. Waiwitlikhit, I. Stoica, Y. Sun, T. Hashimoto, D. Kang · 2024 · 41st International Conference on Machine Learning (ICML 2024)
Supports: ZkAudit protocol; models and datasets; accuracy; costs; assumptions; architecture disclosure; data poisoning
Locator: abstract; §5; Tables 1-4; limitations
Version and catalogue details - S-0009 / Tier B
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments ↗
C. Schnabl, D. Hugenroth, B. Marino, A. R. Beresford · 2025 · ICML 2025 Workshop on Technical AI Governance
Supports: multi-party enclave audit protocol; transparency log; prototype and throughput; CPU versus GPU cost and slowdown; vendor trust
Locator: §3; §4; §5; Table 2
Version and catalogue details - S-1503 / Tier B
Enabling Verifiably-Scoped Monitoring through Large Language Models and Trusted Compute ↗
B. Penchas, G. Zhao, R. Rinberg · 2026 · ICML 2026 Workshop on Technical AI Governance Research
Supports: plan-scoped monitoring protocol
Locator: abstract
Version and catalogue details - S-1504 / Tier C
Auditor-in-a-Box: Tools for Third-Party Auditing ↗
R. Rinberg, B. Penchas · 2026 · LessWrong
Supports: plan definition; reference implementation; process problems; limitations
Locator: whole post
Version and catalogue details - S-1506 / Tier B
Privacy-Preserving AI Verification via Minimal Information Disclosure ↗
S. Abdelghafar, G. Kulp · 2026 · arXiv
Supports: minimal information disclosure framework; one-bit leakage findings; Groth16 variant; limitations
Locator: abstract; introduction; Appendix A; Figure 5; limitations
Version and catalogue details - S-0014 / Tier C
On TEEs for Privacy-Preserving Monitoring in AI Governance ↗
Gloria Z · 2026 · MIRI Technical Governance Team
Supports: completeness and second-CVM gap; vendor root of trust; GPU TEE maturity; multi-GPU inference; treaty threat model
Locator: resource accounting; hardware auditability; physical attack surface
Version and catalogue details - S-3320 / Tier B
Double Blind Evals: Resolving the Dual Confidentiality Dilemma in AI Safety Auditing ↗
A. Trask, S. Messing, V. Pahwa, P. Maham, R. Kolga, A. Frantz, A. Tash, K. Thomas, S. McGregor, G. Balston, P. Paskov, M. Brundage, A. Vij, B. Hillenbrand, A. Karargyris, T. Acosta, J. Fenster, M. Eilish, R. Elasmar, M. Khan, K. van der Veen, R. S, S. Wagh, S. Gabriel, P. Werneck, L. Strahm, K. McDonough, R. Falcon, K. Lum, W. Isaac · 2026 · Google DeepMind
Supports: double-blind evaluation pilot: participants, model, benchmark, GCP Confidential Space on H100 with Intel TDX, PySyft, mutual attestation checks, overhead figure cited from NVIDIA, uninspected model code, Google in the verification path, scaling to many-node clusters
Locator: abstract; architecture; limitations; future work
Version and catalogue details - S-3361 / Tier C
Confidential computing can enable better frontier AI auditing ↗
A. Tlaie Boria · 2026 · Pour Demain
Supports: Pour Demain's interpretability evaluations of GLM-5.1 on Tinfoil (Intel TDX, eight H200 GPUs): enclave-bound tensors, bounded signed exports, overheads, single-session governance
Locator: whole post
Version and catalogue details - S-3321 / Tier B
Confidential Space overview ↗
Google Cloud · 2026 · Google Cloud documentation
Supports: Confidential Space: multi-party roles; data released only to attested workloads; operator has no access; supported TEEs
Locator: overview
Version and catalogue details - S-3322 / Tier B
Confidential Space release notes ↗
Google Cloud · 2026 · Google Cloud documentation
Supports: Confidential Space generally available, including on H100 GPUs from 2026-04-29
Locator: release notes, 2023-03-28 and 2026-04-29
Version and catalogue details - S-3126 / Tier A
DDRop: Active Memory Interposer Attacks on Confidential VMs by Dropping DDR5 Writes ↗
J. De Meulemeester, S. Gloor, P. Jattke, D. Moghimi, D. Oswald, M. Thompson, K. Razavi, I. Verbauwhede, J. Van Bulck · 2026 · 2026 ACM SIGSAC Conference on Computer and Communications Security (CCS '26)
Supports: DDRop forges attestation reports on an up-to-date Intel TDX platform with an active DDR5 interposer
Locator: site summary
Version and catalogue details - S-1202 / Tier A
TEE.fail: Breaking Trusted Execution Environments via DDR5 Memory Bus Interposition ↗
J. Chuang, A. Seto, N. Berrios, S. van Schaik, C. Garman, D. Genkin · 2026 · 2026 IEEE Symposium on Security and Privacy (SP)
Supports: physical extraction of Intel attestation keys and SEV-SNP signing keys; forged attestations against NVIDIA GPU confidential computing; vendor acknowledgement and positions
Locator: project site summary; paper abstract and disclosure
Version and catalogue details - S-1210 / Tier A
Battering RAM: Low-Cost Interposer Attacks on Confidential Computing via Dynamic Memory Aliasing ↗
J. De Meulemeester, D. Oswald, I. Verbauwhede, J. Van Bulck · 2026 · 47th IEEE Symposium on Security and Privacy (S&P 2026)
Supports: Battering RAM forges SEV-SNP attestation with a DDR4 interposer
Version and catalogue details - S-1212 / Tier A
RMPocalypse: How a Catch-22 Breaks AMD SEV-SNP ↗
B. Schlüter, S. Shinde · 2025 · 2025 ACM SIGSAC Conference on Computer and Communications Security (CCS '25)
Supports: RMPocalypse forges SEV-SNP attestation from a malicious hypervisor
Version and catalogue details - S-1213 / Tier B
SEV-SNP RMP Initialization Vulnerability (AMD-SB-3020) ↗
AMD · 2025 · AMD product security bulletin
Supports: AMD firmware fixes for RMPocalypse (CVE-2025-0033)
Version and catalogue details
- Source review date
- 2026-09-25
- Drafted by (source map)
- ai
- Review handles (source map)
- codex-review