01 / The mechanism and its boundary
What is being described
Weight exfiltration is the unauthorized copying of a model's trained parameters, its weights, out of the environment meant to contain them S-1610 S-0015.
RAND researchers identified 38 meaningfully distinct attack vectors for stealing frontier model weights, and defined five security levels for defending against actors ranging from opportunistic criminals to highly resourced nation-state operations S-1610. Exfiltration can be covert. An attacker who controls an inference server could hide weights inside ordinary model responses using steganography S-0015. Rinberg and colleagues verify inference outputs against a reference to limit what responses can carry. On the MoE-Qwen-30B model, under benign prompts, their detector reduced exfiltratable information to under 0.5% at a false-positive rate below 0.01% S-0015, the approach of bounding unexplained information in outputs. Physical side channels offer another route past network monitoring S-0038, the target of side-channel suppression. In verification the concern runs both ways. The verifier's equipment could leak the prover's secrets, so one low-trust design commits model checkpoints to an independent governing body rather than revealing them, and sends only hashes to the verifier outside the facility S-0018.
Connections in the research map
Related research
Sources and provenance
- S-1610 / Tier B
Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models ↗
S. Nevo, D. Lahav, A. Karpur, Y. Bar-On, H. A. Bradley, J. Alstott · 2024 · RAND Corporation
Supports: model weights as a target for theft; 38 attack vectors; 5 security levels; adversaries from opportunistic criminals to nation-state operations
Locator: summary
Version and catalogue details - S-0015 / Tier B
Verifying LLM Inference to Detect Model Weight Exfiltration ↗
R. Rinberg, A. Karvonen, A. Hoover, D. Reuter, K. Warr · 2025 · arXiv
Supports: exfiltration by hiding weights in model responses (steganography); security game; verification against a reference; on MoE-Qwen-30B, exfiltratable information under 0.5% at a false-positive rate under 0.01%
Locator: abstract
Version and catalogue details - S-0038 / Tier C
Suppressing Side Channels in an Untrusted Data Center via Retrofitted Defenses ↗
N. Cankaya · 2026 · MIRI Technical Governance Team
Supports: physical side channels can bypass network monitoring
Locator: side channels of concern
Version and catalogue details - S-0018 / Tier B
A System Overview for Near-Term, Low-Trust AI Compute Verification ↗
N. Cankaya · 2026 · Machine Intelligence Research Institute
Supports: malicious verifier may aim to exfiltrate prover secrets; checkpoints committed, not revealed, to an independent governing body; only hashes sent to a verifier outside the facility
Locator: threat model; weights
Version and catalogue details
- Source review date
- 2026-09-25
- Drafted by (source map)
- ai
- Review handles (source map)
- codex-review