I-0011 / Isolation & architecture

AI 2040 inference-only verification stack

A proposed retrofit that isolates data-centre inference units, taps their front-end traffic and recomputes random samples to check that only declared inference runs.

R1 ProposedSource reviewed 2026-09-25Provider-reported evidence

01 / The mechanism and its boundary

What the technique establishes

The AI 2040 plan proposes checking that retrofitted data centres run only inference, not training or experiments. Data centres would be split into inference units by removing much of their back-end networking. Passive network taps would copy all front-end traffic to a recomputation server, which re-runs random samples of workloads organized into reproducible packets. Strong physical security would protect the equipment, and side-channel limits plus periodic memory wiping are meant to stop hidden work from leaving results. As of July 2026, passive optical taps were reported to work at 400G and recomputation schemes were being tested, but the integrated stack had not been built; one team working on it says nothing is past a proof-of-principle prototype. The obstacles are reproducible inference and networking, securing the recomputation server, and scaling side-channel defences. It does not verify that spare compute is unused. The 400G tap claim concerns components, not an installed system.

Threat model
Adversarial prover
Adversarial evaluation
None
Hardware needed
Retrofit device
Prover cooperation
Required
Confidentiality
Partial
Category
Isolation & architecture
Technical detail and cited results
  • Secure network gateways. The plan names more capable gateways as a possible replacement for the passive taps. Splitters on the input and output paths log data flow into packets, which are randomly sampled and sent to a secure recomputing server. At boot, the gateways would also check that only whitelisted model weights are uploaded to the inference unit S-0067.
  • Physical security measures. The plan lists tamper-evident enclosures, security cameras, perimeter controls and air-gapping S-0067.
  • Other options the plan names. Zero-knowledge proofs could preserve privacy and need less hardware retrofit, but are "currently somewhat speculative" for lack of efficient enough algorithms. Memory-challenge verification might also avoid a hardware retrofit, but "seems likely to face some difficulties with verifying completeness". For an initial phase, removing major scale-out interconnect and installing simple sensors might suffice S-0067.

Claims and scope

A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.

Readiness for a stated use

R1 Proposed

Assessed use: showing that retrofitted data centres run only inference

medium confidence · current · assessed 2026-09-25 · rubric 1.1

This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.

The six-layer design, its claim and its assumptions are published, but the integrated stack has not been built.

  • R1 met: the plan publishes the design of the stack, the claim it would verify (only inference runs on retrofitted compute) and its key assumptions, including a trusted recomputation server and strong physical security S-0067.
  • R2 not met: there is no public implementation or end-to-end result for the integrated stack. The plan's companion page lists a reproducible inference stack as not started and gives no plan yet for scaling side-channel defences S-1511. A team working on the components says that nothing is past a proof-of-principle prototype S-1512. Some components are further along: passive optical taps work at 400G, and recomputation schemes such as TOPLOC and DiFR are being tested on relevant hardware S-1511; see Sampled inference recomputation. Amodo reports a lab prototype that feeds a passive tap on a 1 Gbps link into DiFR recomputation, which covers two of the six layers S-1312.

Confidence is medium: the design is clearly published, but component status comes from pages that are updated over time S-1511 S-1008.

Evidence needed for the next level

  • A public working implementation or reproducible end-to-end results for the integrated stack (isolated inference unit, taps, packetization and recomputation) at realistic scale or against a stated adversary.

  • A hardened recomputation server and a method for checking that taps are correctly installed and remain in place.

  • Red-teaming of recomputation and of the completeness measures (side channels, memory wiping).

Limitations, flaws, and blockers

These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.

significant / open / theoretical argument

The recomputation server must be trusted

The plan calls the integrity of the recomputation server an extremely important aspect, and its argument that sampling verifies all outputs assumes that the server's computations and outputs can be trusted. The companion page notes that the server sits inside the prover's facility, possibly under the prover's physical control, and that hardening it against integrity attacks needs significant research. Amodo rates recomputation-server security as not on track.

S-0067S-1511S-1008

significant / open / theoretical argument

Spare compute is not verified

The plan states that it does not verify that spare compute is unused for unapproved workloads, because this seems very challenging. It relies instead on side-channel bounds and memory wipes, so that the only results that persist are verified inference outputs.

S-0067

significant / open / demonstrated attack

A recomputation family degrades against prompt-controlling adversaries

The plan's companion page names DiFR among the recomputation schemes being tested. An independent study found that Gumbel-based inference verification, the family that includes Token-DiFR, leaks roughly twice as many bits per token when the adversary chooses prompts that disrupt the text's structure, across six models of 1 to 32 billion parameters. The slowdown it imposes on a hidden-signalling adversary falls from 146–254 times under benign prompts to 60–118 times. The attack weakens the bound on hidden information in outputs, which the plan relies on to keep undeclared results from leaving.

S-1511S-0015S-1507S-0067

What still blocks use or stronger assurance

  1. A fully reproducible inference stack needs substantial software and tooling, and per-packet network reproducibility may need considerable software, firmware and possibly hardware work.

    Dependency: Whole-workload recomputation (reproducible packets)

    S-1511
  2. Passive optical taps work at 400G, but the 800G and 1600G line rates now arriving in data centres are undemonstrated.

    Dependency: Network taps and certifiers

    S-1511
  3. Checking that taps are correctly installed and stay in place at scale is not a solved problem, and hardening the recomputation server inside the prover's facility needs significant research.

    Dependency: Tamper evidence for verifier devices

    S-1511S-1008
  4. There is no plan yet for quickly scaling side-channel defences on a frontier cluster; only early theoretical pieces exist.

    Dependency: Side-channel suppression for isolated facilities

    S-1511
  5. Memory wiping may use existing algorithms, but hardware testing is at an early stage.

    Dependency: Memory wiping and proofs of secure erasure

    S-1511
  6. Robust red-teaming of recomputation schemes has not started, and most algorithm development remains academic.

    S-1511S-1008

Connections in the research map

Depends on

Mechanisms implemented

Concepts used

Organizations and developers

Sources and provenance

  1. S-0067 / Tier C

    Verification Plan ↗

    R. Dean · 2026 · AI 2040

    Supports: the six-layer stack; isolation rationale; passive optical taps; packets and reproducibility; recomputation and its trust assumption; physical security measures; completeness measures; spare-compute scope; gateway variant; other options

    Locator: Summary of the plan; inference-only retrofit description; Appendix reference

    Version and catalogue details
  2. S-1511 / Tier C

    Get Involved in Verification ↗

    AI Futures Project · 2026 · AI 2040

    Supports: component status and open problems as of 9 July 2026; recomputation-server hardening

    Locator: network taps; reproducible packets; partial recomputation; physical security; completeness

    Version and catalogue details
  3. S-1512 / Tier C

    Verifying international AI deals: Plan A, the state-of-play, and what you can do to help ↗

    T. Milton, S. Reynolds, C. Jacobi, J. Foster · 2026 · Amodo (Substack)

    Supports: overall maturity statement; description of taps and recomputation

    Locator: introduction

    Version and catalogue details
  4. S-1008 / Tier C

    AI 2040 Plan A — Verification SITREP ↗

    Amodo Design · 2026 · Amodo Design

    Supports: status of recomputation server security and red-teaming

    Locator: status items

    Version and catalogue details
  5. S-1312 / Tier C

    Fitting a Network TAP to our Inference Verification Prototype ↗

    Amodo Design · 2026 · Amodo Design

    Supports: passive tap fitted to a DiFR recomputation prototype on a 1 Gbps lab link; stress-test result; need for an active tap

    Locator: setup, results and side-channel sections

    Version and catalogue details
  6. S-0015 / Tier B

    Verifying LLM Inference to Detect Model Weight Exfiltration ↗

    R. Rinberg, A. Karvonen, A. Hoover, D. Reuter, K. Warr · 2025 · arXiv

    Supports: Token-DiFR as a Gumbel-Max estimator

    Locator: §6.3

    Version and catalogue details
  7. S-1507 / Tier B

    Adversarial Entropy Inflation Against Gumbel-Based Inference Verification ↗

    N. Kezins · 2026 · arXiv

    Supports: prompt-controlling attack on Gumbel-based verification; models tested; bits per token; slowdown factor

    Locator: abstract

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review