01 / The mechanism and its boundary
What is being described
The claim is that no copy of specified model weights has left a designated facility through networks, physical media or other channels. Agreements and security commitments may require this, for example against theft by well-resourced attackers. Verification itself creates a related concern, since monitoring equipment could leak the prover's secrets. It is a negative claim over every channel: network links, physical media, people and physical side channels. Approaches include checking that outgoing traffic is explained by legitimate computation so that hidden data stands out, capping the bandwidth of external links, suppressing side channels, and physical security with inspection. At best they bound how much could leave. Checking inference outputs against a reference cut the information an attacker could hide in responses to under 0.5% in one test; in an independent study, an attacker who controlled the prompts roughly doubled the leak. Physical and insider channels remain harder to bound.
State of verification
Editorial synthesis from the AI Verification Tech Map.
Only the network channel has a demonstrated (R2) bound, and only for inference outputs. Physical side channels rely on proposed (R1) designs, and insiders and physical media are left to facility security S-1610 and personnel measures such as whistleblowing S-0002.
For example, a verifier may want to know that a facility serving a model through an API has not let the weights out inside its responses.
- Network taps (R1) record what the facility sends, and sampled recomputation re-runs sampled requests on a trusted copy of the model with the same random seed S-0015. Bounding unexplained information (R2) then scores how plausible each returned token is under honest sampling, which bounds how fast hidden bits can pass unnoticed S-0015.
- Rinberg and colleagues published code for this check. On a 30-billion-parameter model under benign prompts, it cut the information an attacker could hide to under 0.5%, at a false-positive rate below 0.01% S-0015.
- This establishes a rate limit on one channel, not that no copy left. In an independent study, an attacker who controlled the prompts roughly doubled the leak per token S-1507. One estimate still leaves about 40 Mbit/s of covert output for a 200,000-GPU inference cluster at full load, assuming about 0.1 bits per token survive replay checks S-1300. Deterministic inference would remove the tolerance that hidden data exploits S-0020.
- The bound holds only if checked outputs are the only way out. Caps on external bandwidth (Communication between compute groups is bounded) limit the total that can leave S-1508. Side-channel suppression (R1) is designed to hold physical leakage to a few kilobits per second, not to zero S-0038. RAND's secure inference data centre (R1) is designed to protect weights against state-backed attackers but does not describe how an outside party would verify it S-1510.
Connections in the research map
Techniques addressing this claim
Sources and provenance
- S-1610 / Tier B
Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models ↗
S. Nevo, D. Lahav, A. Karpur, Y. Bar-On, H. A. Bradley, J. Alstott · 2024 · RAND Corporation
Supports: importance of protecting frontier weights; 38 attack vectors; five security levels; adversaries up to nation-states; comprehensive defences
Locator: summary
Version and catalogue details - S-0015 / Tier B
Verifying LLM Inference to Detect Model Weight Exfiltration ↗
R. Rinberg, A. Karvonen, A. Hoover, D. Reuter, K. Warr · 2025 · arXiv
Supports: steganographic exfiltration via inference server responses; security game; fixed-seed re-run and token-plausibility check; code release; <0.5% exfiltratable information at <0.01% FPR on MoE-Qwen-30B under benign prompts; >200x adversary slowdown
Locator: §1 Contributions; §5–7
Version and catalogue details - S-1507 / Tier B
Adversarial Entropy Inflation Against Gumbel-Based Inference Verification ↗
N. Kezins · 2026 · arXiv
Supports: prompt-controlling adversary roughly doubles bits leaked per token; slowdown falls to 60–118x
Locator: abstract; results
Version and catalogue details - S-0005 / Tier B
Mechanisms to Verify International Agreements About AI Development ↗
A. Scher, L. Thiergart · 2025 · arXiv
Supports: strong security to keep weights in a data centre plus close monitoring
Locator: data-centre security discussion
Version and catalogue details - S-0018 / Tier B
A System Overview for Near-Term, Low-Trust AI Compute Verification ↗
N. Cankaya · 2026 · Machine Intelligence Research Institute
Supports: verifier aims to exfiltrate prover secrets; only commitments leave the facility; egress explainable by ingress; memory wiping; side-channel target
Locator: threat model; architecture; open problems
Version and catalogue details - S-0038 / Tier C
Suppressing Side Channels in an Untrusted Data Center via Retrofitted Defenses ↗
N. Cankaya · 2026 · MIRI Technical Governance Team
Supports: physical side channels can bypass network monitoring; defences
Locator: channels of concern; defences
Version and catalogue details - S-0002 / Tier B
Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment ↗
M. Baker, G. Kulp, O. Marks, M. Brundage, L. Heim · 2025 · RAND Corporation
Supports: confidentiality as protecting models, data and code from theft; whistleblower and interview layers
Locator: §1; §4
Version and catalogue details - S-0031 / Tier C
The Fundamentals and Feasibility of Secure Network Taps for Verifying AI Datacenter Use ↗
N. Cankaya · 2026 · The Datacenter Lie Detector
Supports: active taps scrubbing headers against covert channels on front-end links
Locator: frontend vs backend
Version and catalogue details - S-1508 / Tier B
Haiku to Opus in Just 10 bits: LLMs Unlock Large Compression Gains ↗
R. Rinberg, A. M. Carrell, S. Henniger, N. Carlini, K. Warr · 2026 · arXiv
Supports: egress limits cap what can be stolen
Locator: §5.1
Version and catalogue details - S-0020 / Tier B
Bit-Exact AI Inference Verification Without Performance Tradeoffs ↗
N. Cankaya · 2026 · ICML 2026 Workshop on Technical AI Governance Research
Supports: approximate output matching leaves degrees of freedom, including steganography, that covert adversaries can exploit
Locator: abstract
Version and catalogue details - S-1510 / Tier B
Highly Secure Inference Data Centers: A Vertically Integrated Strategy for Security Engineering ↗
S. F. Comer, H. Pavela, V. Gandhi, K. Siler-Evans, E. Devendorf, B. Kelley, J. Gimbi, J. Aguirre, G. Kulp, M. Stalczynski, M. J. Malone · 2026 · RAND Corporation (Research Report RR-A4827-1)
Supports: secure inference data centre design against state-backed attackers; no external verification path described
Locator: Summary; ch. 1
Version and catalogue details - S-1300 / Tier B
Fingerprinting All AI Cluster I/O Without Mutually Trusted Processors ↗
N. Cankaya, J. Kryś, J. Ng, L. Marks, F. Krückel · 2026 · arXiv
Supports: residual covert egress of about 40 Mbit/s for a 200k-GPU inference cluster at about 0.1 bits per token after replay checks
Locator: §5.2
Version and catalogue details
- Source review date
- 2026-09-25
- Drafted by (source map)
- ai
- Review handles (source map)
- codex-review