C-0009

Model weights have not left the facility

No copy of specified model weights has left a designated facility through networks, physical media or other channels.

Source reviewed 2026-09-25

01 / The mechanism and its boundary

What is being described

The claim is that no copy of specified model weights has left a designated facility through networks, physical media or other channels. Agreements and security commitments may require this, for example against theft by well-resourced attackers. Verification itself creates a related concern, since monitoring equipment could leak the prover's secrets. It is a negative claim over every channel: network links, physical media, people and physical side channels. Approaches include checking that outgoing traffic is explained by legitimate computation so that hidden data stands out, capping the bandwidth of external links, suppressing side channels, and physical security with inspection. At best they bound how much could leave. Checking inference outputs against a reference cut the information an attacker could hide in responses to under 0.5% in one test; in an independent study, an attacker who controlled the prompts roughly doubled the leak. Physical and insider channels remain harder to bound.

State of verification

Editorial synthesis from the AI Verification Tech Map.

Only the network channel has a demonstrated (R2) bound, and only for inference outputs. Physical side channels rely on proposed (R1) designs, and insiders and physical media are left to facility security S-1610 and personnel measures such as whistleblowing S-0002.

For example, a verifier may want to know that a facility serving a model through an API has not let the weights out inside its responses.

  • Network taps (R1) record what the facility sends, and sampled recomputation re-runs sampled requests on a trusted copy of the model with the same random seed S-0015. Bounding unexplained information (R2) then scores how plausible each returned token is under honest sampling, which bounds how fast hidden bits can pass unnoticed S-0015.
  • Rinberg and colleagues published code for this check. On a 30-billion-parameter model under benign prompts, it cut the information an attacker could hide to under 0.5%, at a false-positive rate below 0.01% S-0015.
  • This establishes a rate limit on one channel, not that no copy left. In an independent study, an attacker who controlled the prompts roughly doubled the leak per token S-1507. One estimate still leaves about 40 Mbit/s of covert output for a 200,000-GPU inference cluster at full load, assuming about 0.1 bits per token survive replay checks S-1300. Deterministic inference would remove the tolerance that hidden data exploits S-0020.
  • The bound holds only if checked outputs are the only way out. Caps on external bandwidth (Communication between compute groups is bounded) limit the total that can leave S-1508. Side-channel suppression (R1) is designed to hold physical leakage to a few kilobits per second, not to zero S-0038. RAND's secure inference data centre (R1) is designed to protect weights against state-backed attackers but does not describe how an outside party would verify it S-1510.

Connections in the research map

Concepts used

Techniques addressing this claim

Sources and provenance

  1. S-1610 / Tier B

    Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models ↗

    S. Nevo, D. Lahav, A. Karpur, Y. Bar-On, H. A. Bradley, J. Alstott · 2024 · RAND Corporation

    Supports: importance of protecting frontier weights; 38 attack vectors; five security levels; adversaries up to nation-states; comprehensive defences

    Locator: summary

    Version and catalogue details
  2. S-0015 / Tier B

    Verifying LLM Inference to Detect Model Weight Exfiltration ↗

    R. Rinberg, A. Karvonen, A. Hoover, D. Reuter, K. Warr · 2025 · arXiv

    Supports: steganographic exfiltration via inference server responses; security game; fixed-seed re-run and token-plausibility check; code release; <0.5% exfiltratable information at <0.01% FPR on MoE-Qwen-30B under benign prompts; >200x adversary slowdown

    Locator: §1 Contributions; §5–7

    Version and catalogue details
  3. S-1507 / Tier B

    Adversarial Entropy Inflation Against Gumbel-Based Inference Verification ↗

    N. Kezins · 2026 · arXiv

    Supports: prompt-controlling adversary roughly doubles bits leaked per token; slowdown falls to 60–118x

    Locator: abstract; results

    Version and catalogue details
  4. S-0005 / Tier B

    Mechanisms to Verify International Agreements About AI Development ↗

    A. Scher, L. Thiergart · 2025 · arXiv

    Supports: strong security to keep weights in a data centre plus close monitoring

    Locator: data-centre security discussion

    Version and catalogue details
  5. S-0018 / Tier B

    A System Overview for Near-Term, Low-Trust AI Compute Verification ↗

    N. Cankaya · 2026 · Machine Intelligence Research Institute

    Supports: verifier aims to exfiltrate prover secrets; only commitments leave the facility; egress explainable by ingress; memory wiping; side-channel target

    Locator: threat model; architecture; open problems

    Version and catalogue details
  6. S-0038 / Tier C

    Suppressing Side Channels in an Untrusted Data Center via Retrofitted Defenses ↗

    N. Cankaya · 2026 · MIRI Technical Governance Team

    Supports: physical side channels can bypass network monitoring; defences

    Locator: channels of concern; defences

    Version and catalogue details
  7. S-0002 / Tier B

    Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment ↗

    M. Baker, G. Kulp, O. Marks, M. Brundage, L. Heim · 2025 · RAND Corporation

    Supports: confidentiality as protecting models, data and code from theft; whistleblower and interview layers

    Locator: §1; §4

    Version and catalogue details
  8. S-0031 / Tier C

    The Fundamentals and Feasibility of Secure Network Taps for Verifying AI Datacenter Use ↗

    N. Cankaya · 2026 · The Datacenter Lie Detector

    Supports: active taps scrubbing headers against covert channels on front-end links

    Locator: frontend vs backend

    Version and catalogue details
  9. S-1508 / Tier B

    Haiku to Opus in Just 10 bits: LLMs Unlock Large Compression Gains ↗

    R. Rinberg, A. M. Carrell, S. Henniger, N. Carlini, K. Warr · 2026 · arXiv

    Supports: egress limits cap what can be stolen

    Locator: §5.1

    Version and catalogue details
  10. S-0020 / Tier B

    Bit-Exact AI Inference Verification Without Performance Tradeoffs ↗

    N. Cankaya · 2026 · ICML 2026 Workshop on Technical AI Governance Research

    Supports: approximate output matching leaves degrees of freedom, including steganography, that covert adversaries can exploit

    Locator: abstract

    Version and catalogue details
  11. S-1510 / Tier B

    Highly Secure Inference Data Centers: A Vertically Integrated Strategy for Security Engineering ↗

    S. F. Comer, H. Pavela, V. Gandhi, K. Siler-Evans, E. Devendorf, B. Kelley, J. Gimbi, J. Aguirre, G. Kulp, M. Stalczynski, M. J. Malone · 2026 · RAND Corporation (Research Report RR-A4827-1)

    Supports: secure inference data centre design against state-backed attackers; no external verification path described

    Locator: Summary; ch. 1

    Version and catalogue details
  12. S-1300 / Tier B

    Fingerprinting All AI Cluster I/O Without Mutually Trusted Processors ↗

    N. Cankaya, J. Kryś, J. Ng, L. Marks, F. Krückel · 2026 · arXiv

    Supports: residual covert egress of about 40 Mbit/s for a 200k-GPU inference cluster at about 0.1 bits per token after replay checks

    Locator: §5.2

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review