01 / The mechanism and its boundary
What the technique establishes
OpenMined's PySyft coordinates evaluations in which a model owner keeps its weights from the evaluator and the evaluator keeps its prompts from the model owner. Both parties check an enclave's attestation, submit code and assets, and approve the code before it runs. In 2026, AVERI evaluated Gemini 2.5 Flash Lite on private MLCommons prompts in Google Cloud Confidential Space using PySyft v0.10.x on an NVIDIA H100 with Intel TDX. A separate evaluation with Singapore AISI used a private prompt set. The participants report end-to-end operation, but not all model code could be inspected or allowlisted. The guest operating system's builds were not independently reproducible, and Google's services signed and verified the attestation. The participant report contains no independent security evaluation of this workflow.
- Threat model
- Semi-trusted prover
- Adversarial evaluation
- Published analysis
- Hardware needed
- Existing hardware features
- Prover cooperation
- Required
- Confidentiality
- Preserving
- Category
- Cryptography & computation
Claims and scope
A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.
The declared evaluation was run (upstream draft)
Both parties approve an attested evaluation on their submitted assets. The pilot accepted proprietary model code that could not all be inspected or allowlisted; it does not establish ongoing serving identity (S-3320).
Readiness for a stated use
Assessed use: evaluating a private model on private prompts, neither party seeing the other's inputs
medium confidence · current · assessed 2026-09-25 · rubric 1.1
This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.
A participant report describes an end-to-end evaluation with private assets on commercial GPU hardware; PySyft's double-blind workflow has not been shown as a generally available service.
- R1 met: the report states the mutual-confidentiality claim, the enclave trust assumptions and the submission and approval procedure S-3320.
- R2 met: AVERI evaluated Gemini 2.5 Flash Lite with private MLCommons prompts on an H100 with Intel TDX and PySyft v0.10.x. The report gives the stack and workflow, and reports a separate evaluation with Singapore AISI that used private prompts S-3320.
- R3 not met for this implementation: OpenMined documents a pilot with real private assets, but not an available production service or reliance on its result for a verification decision S-3320 S-3563.
- R4 not met: the participant report includes no independent public security evaluation of the workflow. Confidence is medium. The demonstration and its limits come from the participants' own report S-3320.
Evidence needed for the next level
A generally available production workflow, or documented reliance by another party on its result for a verification decision.
An independent public security evaluation that leaves no critical flaw open.
Limitations, flaws, and blockers
These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.
significant / open / open question
Physical attack boundary
The TEE findings include physical-host attacks that forge TDX attestations and a demonstration pairing forged TDX evidence with relayed H100 attestations S-1202 S-3126. The pilot's report describes a TDX and H100 deployment, but does not test that deployment against these attacks S-3320.
What still blocks use or stronger assurance
The pilot could not inspect or allowlist all model code, and the guest operating system builds were not independently reproducible.
Dependency: TEE remote attestation for AI workloads
S-3320- S-3320
The pilot ran on one H100; the authors name many-node confidential GPU clusters as the next scale target.
Connections in the research map
Depends on
- TEE remote attestation for AI workloads
Both parties rely on remote attestation of the Intel TDX and NVIDIA H100 confidential-computing stack.
Mechanisms implemented
Organizations and developers
Sources and provenance
- S-3320 / Tier B
Double Blind Evals: Resolving the Dual Confidentiality Dilemma in AI Safety Auditing ↗
A. Trask, S. Messing, V. Pahwa, P. Maham, R. Kolga, A. Frantz, A. Tash, K. Thomas, S. McGregor, G. Balston, P. Paskov, M. Brundage, A. Vij, B. Hillenbrand, A. Karargyris, T. Acosta, J. Fenster, M. Eilish, R. Elasmar, M. Khan, K. van der Veen, R. S, S. Wagh, S. Gabriel, P. Werneck, L. Strahm, K. McDonough, R. Falcon, K. Lum, W. Isaac · 2026 · Google DeepMind
Supports: procedure, trust boundary, participants, model, private prompts, hardware, results and limits
Locator: §2.5; §3; §4
Version and catalogue details - S-3563 / Tier C
PySyft used for first double-blind evaluation of a proprietary, frontier-class AI model ↗
OpenMined Team · 2026 · OpenMined
Supports: OpenMined's description of PySyft and the two 2026 evaluations
Locator: Executive Summary
Version and catalogue details - S-1202 / Tier A
TEE.fail: Breaking Trusted Execution Environments via DDR5 Memory Bus Interposition ↗
J. Chuang, A. Seto, N. Berrios, S. van Schaik, C. Garman, D. Genkin · 2026 · 2026 IEEE Symposium on Security and Privacy (SP)
Supports: inherited Intel TDX physical-host attestation forgery and H100 relay; not a PySyft workflow evaluation
Locator: Abstract; §8.3; site FAQ
Version and catalogue details - S-3126 / Tier A
DDRop: Active Memory Interposer Attacks on Confidential VMs by Dropping DDR5 Writes ↗
J. De Meulemeester, S. Gloor, P. Jattke, D. Moghimi, D. Oswald, M. Thompson, K. Razavi, I. Verbauwhede, J. Van Bulck · 2026 · 2026 ACM SIGSAC Conference on Computer and Communications Security (CCS '26)
Supports: inherited DDR5 physical-host attack on Intel TDX; not a PySyft workflow evaluation
Locator: Threat model; TDX case studies
Version and catalogue details
- Source review date
- 2026-09-25
- Drafted by (source map)
- ai
- Review handles (source map)
- codex-review