K-0022

Weight exfiltration

Unauthorized copying of a model's trained parameters out of the environment meant to contain them, by theft or through covert channels.

Source reviewed 2026-09-25

01 / The mechanism and its boundary

What is being described

Weight exfiltration is the unauthorized copying of a model's trained parameters, its weights, out of the environment meant to contain them S-1610 S-0015.

RAND researchers identified 38 meaningfully distinct attack vectors for stealing frontier model weights, and defined five security levels for defending against actors ranging from opportunistic criminals to highly resourced nation-state operations S-1610. Exfiltration can be covert. An attacker who controls an inference server could hide weights inside ordinary model responses using steganography S-0015. Rinberg and colleagues verify inference outputs against a reference to limit what responses can carry. On the MoE-Qwen-30B model, under benign prompts, their detector reduced exfiltratable information to under 0.5% at a false-positive rate below 0.01% S-0015, the approach of bounding unexplained information in outputs. Physical side channels offer another route past network monitoring S-0038, the target of side-channel suppression. In verification the concern runs both ways. The verifier's equipment could leak the prover's secrets, so one low-trust design commits model checkpoints to an independent governing body rather than revealing them, and sends only hashes to the verifier outside the facility S-0018.

Connections in the research map

Related research

Sources and provenance

  1. S-1610 / Tier B

    Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models ↗

    S. Nevo, D. Lahav, A. Karpur, Y. Bar-On, H. A. Bradley, J. Alstott · 2024 · RAND Corporation

    Supports: model weights as a target for theft; 38 attack vectors; 5 security levels; adversaries from opportunistic criminals to nation-state operations

    Locator: summary

    Version and catalogue details
  2. S-0015 / Tier B

    Verifying LLM Inference to Detect Model Weight Exfiltration ↗

    R. Rinberg, A. Karvonen, A. Hoover, D. Reuter, K. Warr · 2025 · arXiv

    Supports: exfiltration by hiding weights in model responses (steganography); security game; verification against a reference; on MoE-Qwen-30B, exfiltratable information under 0.5% at a false-positive rate under 0.01%

    Locator: abstract

    Version and catalogue details
  3. S-0038 / Tier C

    Suppressing Side Channels in an Untrusted Data Center via Retrofitted Defenses ↗

    N. Cankaya · 2026 · MIRI Technical Governance Team

    Supports: physical side channels can bypass network monitoring

    Locator: side channels of concern

    Version and catalogue details
  4. S-0018 / Tier B

    A System Overview for Near-Term, Low-Trust AI Compute Verification ↗

    N. Cankaya · 2026 · Machine Intelligence Research Institute

    Supports: malicious verifier may aim to exfiltrate prover secrets; checkpoints committed, not revealed, to an independent governing body; only hashes sent to a verifier outside the facility

    Locator: threat model; weights

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review