I-0017 / Cryptography & computation

PySyft double-blind evaluations

PySyft coordinates an attested enclave where a model owner and evaluator run tests without sharing weights or private prompts.

R2 DemonstratedSource reviewed 2026-09-25Provider-reported evidence

01 / The mechanism and its boundary

What the technique establishes

OpenMined's PySyft coordinates evaluations in which a model owner keeps its weights from the evaluator and the evaluator keeps its prompts from the model owner. Both parties check an enclave's attestation, submit code and assets, and approve the code before it runs. In 2026, AVERI evaluated Gemini 2.5 Flash Lite on private MLCommons prompts in Google Cloud Confidential Space using PySyft v0.10.x on an NVIDIA H100 with Intel TDX. A separate evaluation with Singapore AISI used a private prompt set. The participants report end-to-end operation, but not all model code could be inspected or allowlisted. The guest operating system's builds were not independently reproducible, and Google's services signed and verified the attestation. The participant report contains no independent security evaluation of this workflow.

Threat model
Semi-trusted prover
Adversarial evaluation
Published analysis
Hardware needed
Existing hardware features
Prover cooperation
Required
Confidentiality
Preserving
Category
Cryptography & computation

Claims and scope

A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.

Readiness for a stated use

R2 Demonstrated

Assessed use: evaluating a private model on private prompts, neither party seeing the other's inputs

medium confidence · current · assessed 2026-09-25 · rubric 1.1

This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.

A participant report describes an end-to-end evaluation with private assets on commercial GPU hardware; PySyft's double-blind workflow has not been shown as a generally available service.

  • R1 met: the report states the mutual-confidentiality claim, the enclave trust assumptions and the submission and approval procedure S-3320.
  • R2 met: AVERI evaluated Gemini 2.5 Flash Lite with private MLCommons prompts on an H100 with Intel TDX and PySyft v0.10.x. The report gives the stack and workflow, and reports a separate evaluation with Singapore AISI that used private prompts S-3320.
  • R3 not met for this implementation: OpenMined documents a pilot with real private assets, but not an available production service or reliance on its result for a verification decision S-3320 S-3563.
  • R4 not met: the participant report includes no independent public security evaluation of the workflow. Confidence is medium. The demonstration and its limits come from the participants' own report S-3320.

Evidence needed for the next level

  • A generally available production workflow, or documented reliance by another party on its result for a verification decision.

  • An independent public security evaluation that leaves no critical flaw open.

Limitations, flaws, and blockers

These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.

significant / open / open question

Physical attack boundary

The TEE findings include physical-host attacks that forge TDX attestations and a demonstration pairing forged TDX evidence with relayed H100 attestations S-1202 S-3126. The pilot's report describes a TDX and H100 deployment, but does not test that deployment against these attacks S-3320.

S-3320S-1202S-3126
Response recorded by the source map

The attack researchers report that physical interposer attacks are outside Intel's threat model. They recommend physically secure servers S-1202 S-3126.

What still blocks use or stronger assurance

  1. The pilot could not inspect or allowlist all model code, and the guest operating system builds were not independently reproducible.

    Dependency: TEE remote attestation for AI workloads

    S-3320
  2. The pilot ran on one H100; the authors name many-node confidential GPU clusters as the next scale target.

    S-3320

Connections in the research map

Depends on

Mechanisms implemented

Concepts used

Organizations and developers

Sources and provenance

  1. S-3320 / Tier B

    Double Blind Evals: Resolving the Dual Confidentiality Dilemma in AI Safety Auditing ↗

    A. Trask, S. Messing, V. Pahwa, P. Maham, R. Kolga, A. Frantz, A. Tash, K. Thomas, S. McGregor, G. Balston, P. Paskov, M. Brundage, A. Vij, B. Hillenbrand, A. Karargyris, T. Acosta, J. Fenster, M. Eilish, R. Elasmar, M. Khan, K. van der Veen, R. S, S. Wagh, S. Gabriel, P. Werneck, L. Strahm, K. McDonough, R. Falcon, K. Lum, W. Isaac · 2026 · Google DeepMind

    Supports: procedure, trust boundary, participants, model, private prompts, hardware, results and limits

    Locator: §2.5; §3; §4

    Version and catalogue details
  2. S-3563 / Tier C

    PySyft used for first double-blind evaluation of a proprietary, frontier-class AI model ↗

    OpenMined Team · 2026 · OpenMined

    Supports: OpenMined's description of PySyft and the two 2026 evaluations

    Locator: Executive Summary

    Version and catalogue details
  3. S-1202 / Tier A

    TEE.fail: Breaking Trusted Execution Environments via DDR5 Memory Bus Interposition ↗

    J. Chuang, A. Seto, N. Berrios, S. van Schaik, C. Garman, D. Genkin · 2026 · 2026 IEEE Symposium on Security and Privacy (SP)

    Supports: inherited Intel TDX physical-host attestation forgery and H100 relay; not a PySyft workflow evaluation

    Locator: Abstract; §8.3; site FAQ

    Version and catalogue details
  4. S-3126 / Tier A

    DDRop: Active Memory Interposer Attacks on Confidential VMs by Dropping DDR5 Writes ↗

    J. De Meulemeester, S. Gloor, P. Jattke, D. Moghimi, D. Oswald, M. Thompson, K. Razavi, I. Verbauwhede, J. Van Bulck · 2026 · 2026 ACM SIGSAC Conference on Computer and Communications Security (CCS '26)

    Supports: inherited DDR5 physical-host attack on Intel TDX; not a PySyft workflow evaluation

    Locator: Threat model; TDX case studies

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review