I-0012 / Isolation & architecture

Low-trust AI compute verification system overview

A retrofittable reference design in which network taps commit to all facility traffic, and air-gapped, independently sourced checkers later re-run randomly challenged records.

R1 ProposedSource reviewed 2026-09-25Provider-reported evidence

01 / The mechanism and its boundary

What the technique establishes

Naci Cankaya's system overview, published by the Machine Intelligence Research Institute's Technical Governance Team, proposes a privacy-preserving verification system that can be added to existing AI data centres, for an agreement between states that distrust each other. Network taps hash all traffic crossing the monitored boundary, and only these hashes are shared with the verifier. Later, the verifier challenges random records. The operator supplies the matching data, and at least three independently sourced computing clusters, each air-gapped inside the operator's data centre, re-run the workload and screen it, for example for training versus inference or banned uses. One-way optical links, cross-checks of every output and re-transmission by the operator limit what either side can leak. The document is a working draft, and key parts are open research questions. The main obstacles are bit-exact replay, physical security and side-channel defence. The author says the threat model is under-developed.

Threat model
Adversarial prover
Adversarial evaluation
Published analysis
Hardware needed
Retrofit device
Prover cooperation
Required
Confidentiality
Preserving
Category
Isolation & architecture
Technical detail and cited results
  • Capture. Evidence capture runs online at line rate. Taps parse only link-to-transport-layer structure, hash payloads over groups of packets without regard to object boundaries, and emit no timestamps, because trailing bits could encode information S-0018.
  • Replay metadata. Exact replay needs the hardware SKU, the exact weights and quantization format, the parallelism topology, the software versions and custom kernels, and the batch size of each forward pass S-0018.
  • Instructor and executor. An untrusted but fast executor re-runs the workload without seeing the claimed output, and a small, formally verifiable instructor compares its result with the prover's claim. The cited precedent is an FPGA sentry of 3,543 lines of hardware-description code S-0018.
  • Sampling statistics. P(detect at least one) = 1 - (1 - p)^n. Checking 3,000 random samples detects a 0.1% flaw rate with 95.03% probability S-0018.

Claims and scope

A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.

Readiness for a stated use

R1 Proposed

Assessed use: screening challenged records to show declared inference compute is not training

medium confidence · current · assessed 2026-09-25 · rubric 1.1

This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.

A detailed end-to-end design states the claim and threat model, but there is no integrated implementation and there are no results.

  • R1 met: the overview describes the system end to end, the rules it would support (such as inference versus training, model whitelists and blacklisted uses), a worst-case threat model in which both prover and verifier are hostile nation-states, and its practical requirements S-0018.
  • R2 not met. The document is a working draft that sets out the design and open research questions, not results from an integrated system S-0018. Its companion preprint specifies the tap subsystem and states that empirical validation is still required S-1300. Several building blocks remain open, including passive optical splitting at 53–112 GBaud S-0018.

Confidence is medium: the design is detailed, but the author states that its threat model is under-developed S-0018.

Evidence needed for the next level

  • A public working implementation or reproducible end-to-end results for the capture-then-challenge pipeline, at realistic line rates or against a stated adversary.

  • Demonstrated bit-exact replay of production inference inside a secure auditing environment built from independently sourced components.

  • A developed threat model and red-teaming of the side-channel, egress and inspector-agent components.

Limitations, flaws, and blockers

These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.

significant / open / theoretical argument

Mismatches cannot be attributed to cheating or error

A failed hash or replay does not show whether it came from an evasion attempt, a random bit flip or an evaluation error. The author notes that if detected anomalies can plausibly be waved off as malfunctions, deterrence becomes less effective, so the parties need an agreed escalation procedure that ends in attribution.

S-0018

minor / open / theoretical argument

Deliberate faults leak a bit each

A malicious device can leak one bit by deliberately outputting a wrong result, which blocks a disclosure when the cross-comparison fails. The design therefore needs a pre-agreed budget of tolerated faults.

S-0018

significant / open / open question

Inspector agents may be manipulable

Automated compliance screening with LLM-based inspector agents must resist prompt-injection attacks. Adversarially trained systems might hide malicious actions with steganography, which makes backdoor detection an open problem.

S-0018

What still blocks use or stronger assurance

  1. Empirical feasibility of passive optical splitting at 53–112 GBaud under realistic conditions is an open question.

    Dependency: Network taps and certifiers

    S-0018
  2. Exact replay needs complete hardware and software metadata, and the tolerable slowdown from emulation is an open question.

    Dependency: Deterministic and bit-exact inference

    S-0018
  3. Tamper-evident, rapidly mass-manufacturable and retrofittable enclosures for side-channel defence are an open research question, and physical security against covert communication in every monitored data centre is challenging.

    Dependency: Tamper evidence for verifier devices

    S-0018
  4. A mass-manufacturable, good-enough side-channel defence, particularly power-line filtering, has not been constructed or red-teamed.

    Dependency: Side-channel suppression for isolated facilities

    S-0018
  5. Distinguishing one server's DRAM contents from another's by challenge-response timing, and a general challenge-response protocol for diverse data types, are open.

    Dependency: Timed challenge-response and memory-occupation challenges

    S-0018
  6. The threat model is under-developed and needs input from cybersecurity and AI threat-modelling experts.

    S-0018

Connections in the research map

Depends on

Mechanisms implemented

Concepts used

Organizations and developers

Sources and provenance

  1. S-0018 / Tier B

    A System Overview for Near-Term, Low-Trust AI Compute Verification ↗

    N. Cankaya · 2026 · Machine Intelligence Research Institute

    Supports: purpose; rules; threat model; requirements; execution trace; subsystems; engineering approaches; open problems; prior work; sampling statistics

    Locator: §1; §2a-2c; §3.1-3.2; §4.1-4.3; §5.1-5.3; Appendix A1

    Version and catalogue details
  2. S-1300 / Tier B

    Fingerprinting All AI Cluster I/O Without Mutually Trusted Processors ↗

    N. Cankaya, J. Kryś, J. Ng, L. Marks, F. Krückel · 2026 · arXiv

    Supports: companion secure gateway and tap design; demonstration cost; validation and red-teaming still required

    Locator: abstract; discussion of next steps

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review