M-0014 / Isolation & architecture

Bandwidth limits and compartmentalization

Capping or removing network links between groups of accelerators, so that serving within each group still works but large training across groups becomes far slower.

R2 DemonstratedSource reviewed 2026-09-25

01 / The mechanism and its boundary

What the technique establishes

Bandwidth limits cap or remove the links between groups of AI accelerators ("pods"), so that serving models within each pod keeps working but large training runs across pods become slow and costly. Lucid Computing's design caps each 72-GPU pod at 1 Gbps per direction. Lucid's modelled estimates put covert frontier training at about 350 times less efficient, or about 140 times if two unproven techniques both succeed. Both figures assume an auditor randomizes how pods connect to routers; with operator-chosen routing and compression, the inefficiency can fall to about 25 times. A small software prototype shows the traffic gap between training and inference, but as of September 2026 no cap that a verifier can check has been built or red-teamed. The main obstacles are showing that all traffic crosses the capped links and that the enforcing devices can be trusted. The largest known weakness is progress in low-communication training.

Threat model
Adversarial prover
Adversarial evaluation
Published analysis
Hardware needed
Retrofit device
Prover cooperation
Required
Confidentiality
Preserving
Category
Isolation & architecture
Technical detail and cited results
  • Pod and cap. Lucid Computing's Traffic Shaping design models a pod as one NVIDIA GB200 NVL72 rack (72 GPUs). A shaper at each pod's north-south uplink enforces a recommended cap of β = 1 Gbps (125 MB/s) per direction per pod, averaged over 60 seconds, about 14 Mbps per GPU S-1301.
  • Adversary. The reference served model is a "Fable-sized model (estimated 6T parameter MoE)"; the adversary's target is about 10x larger (60T parameters). Strategies analysed: centralized training adapted to the cap, DiLoCo-family low-communication data parallelism (including streaming and quantized variants), and SWARM-family decentralized pipeline parallelism. The analysis grants the adversary perfect overlap of communication and computation, zero pipeline bubbles, zero straggler and fault overhead, free coordination compute outside the pods, aggressive low-precision recipes and the best published communication-reduction techniques S-1301.
  • Results. The headline is a lower bound on the inefficiency multiplier of about 350x, with central estimates of 1,000x–10,000x depending on training architecture. For SWARM-style pipelines the floor falls to about 140x only if two further unproven techniques both succeed (8x lossless activation compression at frontier scale and an unusually narrow 60T-parameter architecture). The appendix summary claims at least 100x, with roughly 1.4–3.5x of margin, after stacking every charitable and speculative assumption S-1301.
  • Conditions. The auditor randomizes the assignment of pods to routers and re-randomizes it periodically, and per-pod local storage is declared, capped and inspected. If the operator could instead assign whole cells of 100 or more pods to one pipeline stage, the bound would collapse to about 90–220x uncompressed and as low as about 25x with compression S-1301.
  • Amodo's node-level limiter. NVIDIA BlueField-3 DPUs replace the NICs on a 400G Ethernet link, reaching 390 Gbps unencrypted and 193 Gbps encrypted between two DPUs S-1313.

Claims and scope

A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.

Readiness for a stated use

R2 Demonstrated

Assessed use: monitoring inter-node traffic with operator-run software on four GPUs

low confidence · current · assessed 2026-10-08 · rubric 1.1

This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.

R2, narrowly: a public software prototype on data-centre GPUs monitors inter-node traffic against a limit and detects low-communication training, but no cap that a verifier can check has been built.

  • R1 met: the AI 2040 plan proposes isolated inference units created by removing back-end networking, on the grounds that high-bandwidth links are mostly needed only for training S-0067. Lucid Computing gives a detailed design with a cap, a pod definition, adversary strategies, favourable assumptions and residual risks S-1301. The MIRI overview describes perimeters of monitored links sized to pods for inference S-0018.
  • R2 met: MIRI's Technical Governance Team published an end-to-end prototype, with code, that monitors inter-node traffic against a threshold on two nodes with four A100 GPUs S-3220. It tested violating training, compliant inference and a stated adversary, DiLoCo training S-3220. The prototype alerts rather than throttles, and the authors call its software-only design "trivially spoofable" S-3220. Amodo reports building a DPU-based per-node rate limiter on 400G links, for weight security under a trusted controller S-1313. Lucid states that its pod-level cap is "still at the design stage and not yet implemented or red-teamed" S-1301. The mechanism's implementations, RAND secure inference data center (SIDC) design and AI 2040 inference-only verification stack, are proposed architectures at R1.
  • R3 not met: no production-grade cap or monitor for this use is publicly available, and no reliance on one by a party other than its developer is documented.

Confidence is low: the R2 evidence is one small, software-only prototype described in a blog post, and it monitors a limit without enforcing one.

Evidence needed for the next level

  • A production-grade cap or bandwidth monitor for this use, or reliance on one by a party other than its developer for a verification decision.

  • Enforcement outside the prover's control, such as a shaper at each pod uplink, built and tested on data-centre hardware (Lucid reports a proof of concept in development).

  • Measured, not only modelled, training inefficiency under a cap at realistic scale, including low-communication methods.

  • Red-teaming of cap bypass (for example through parallel scale-up switches) and of trust in the shaping devices.

Limitations, flaws, and blockers

These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.

significant / open / theoretical argument

Low-communication training reduces the bandwidth training needs

DiLoCo matched fully synchronous training on 8 workers while communicating 500 times less. Rahman writes that this family of methods theoretically allows large-scale training with less than 100 Mbps. Lucid includes these methods in its bounds, but notes that extreme activation compression, architectures with unusually small inter-layer widths, or modular paradigms could erode the margin.

S-1314S-0060S-1301

significant / open / theoretical argument

Operator control of pod routing collapses the bound

Lucid's analysis finds that if the operator can freely assign pods to routers, it could dedicate a whole cell of 100 or more pods to one pipeline stage. The bound then falls to about 90–220x uncompressed and as low as about 25x with compression. The proposed mitigation, auditor-controlled random assignment that is periodically re-randomized, has not been implemented.

S-1301

significant / open / theoretical argument

Undeclared local storage raises per-pod capacity

More memory or storage per pod helps an adversary. Lucid requires per-pod storage to be declared, capped and physically inspected.

S-1301

significant / open / open question

Training within one pod is not covered

Lucid's bounds concern pre-training models larger than the pods are sized for. Training models that fit in one pod, fine-tuning and reinforcement-learning post-training within one pod are outside the modelled threat.

S-1301

significant / open / theoretical argument

Parallel scale-up switches are hard enforcement points

In GB200 topologies, GPUs reach GPUs in other nodes through NVSwitches without a NIC on the path. Amodo notes that limits are hard to enforce there because many switches work in parallel, so compromising one or two would bypass the limit.

S-1313

What still blocks use or stronger assurance

  1. No cap that a verifier can check has been implemented or red-teamed.

    S-1301
  2. The verifier must know that all traffic leaving a pod crosses the capped, monitored links.

    Dependency: Network taps and certifiers

    S-0018
  3. Shaping devices and routing assignments must be trusted by both parties; Amodo has not yet fully analysed resilience to a compromised DPU.

    Dependency: Tamper evidence for verifier devices

    S-1301S-1313
  4. Advances in low-communication training could shrink the margin that the cap enforces.

    S-1314S-0060S-1301

Connections in the research map

Depends on

Complementary techniques

Concepts used

Organizations and developers

Implementations

The Consortium’s case files

Related editorial reviews use the Consortium’s own descriptive scores and review dates. Their scores are separate from the atlas readiness rubric.

FS-03 / ControlThe approved friend groupRead case file ↗

Sources and provenance

  1. S-0067 / Tier C

    Verification Plan ↗

    R. Dean · 2026 · AI 2040

    Supports: isolated inference units by removing back-end networking; rationale; retrofit illustration

    Locator: inference-only retrofitting proposal

    Version and catalogue details
  2. S-1301 / Tier C

    Traffic Shaping for Workload Classification ↗

    Lucid Computing · 2026 · Lucid Computing (Substack)

    Supports: Traffic Shaping design, cap, pod model, adversary strategies and assumptions, results and their conditions, residual risks, status

    Locator: summary; main text; appendices A.7–A.9

    Version and catalogue details
  3. S-1508 / Tier B

    Haiku to Opus in Just 10 bits: LLMs Unlock Large Compression Gains ↗

    R. Rinberg, A. M. Carrell, S. Henniger, N. Carlini, K. Warr · 2026 · arXiv

    Supports: egress limits cap how much can be stolen

    Locator: §5.1

    Version and catalogue details
  4. S-0018 / Tier B

    A System Overview for Near-Term, Low-Trust AI Compute Verification ↗

    N. Cankaya · 2026 · Machine Intelligence Research Institute

    Supports: perimeter of monitored links; perimeter size for inference and training

    Locator: §5.1.1

    Version and catalogue details
  5. S-3220 / Tier C

    De-risking Interconnect Limits for AI Verification ↗

    A. Scher, D. Sarbakysh, A. Moskvin · 2026 · MIRI Technical Governance Team

    Supports: interconnect-limit monitoring prototype, workloads, measured traffic, DiLoCo detection, alerting without throttling, spoofability, longer training, LoRA and RL

    Locator: whole post

    Version and catalogue details
  6. S-3221 / Tier C

    Activating AI Safety Level 3 Protections ↗

    Anthropic · 2025 · Anthropic

    Supports: Anthropic's egress bandwidth controls for model weights

    Locator: security section

    Version and catalogue details
  7. S-1313 / Tier C

    The Tray as a Bandwidth Boundary ↗

    Amodo Design · 2026 · Amodo Design

    Supports: DPU-based node bandwidth boundary, throughput, NVSwitch limitation, compromised-DPU caveat, security purpose

    Locator: whole note

    Version and catalogue details
  8. S-1314 / Tier B

    DiLoCo: Distributed Low-Communication Training of Language Models ↗

    A. Douillard, Q. Feng, A. A. Rusu, R. Chhaparia, Y. Donchev, A. Kuncoro, M. Ranzato, A. Szlam, J. Shen · 2024 · ICML 2024 Workshop on Advancing Neural Network Training (WANT)

    Supports: 500x less communication on 8 workers

    Locator: abstract

    Version and catalogue details
  9. S-0060 / Tier B

    Does Distributed Training Undermine Compute Governance? ↗

    R. Rahman · 2026 · ICML 2026 Workshop on Technical AI Governance Research

    Supports: DiLoCo-family bandwidth needs; bandwidth caps on evaders judged infeasible

    Locator: §1; appendix G.1

    Version and catalogue details
  10. S-0026 / Tier B

    Verifiable constraints on frontier training via proofs of compartmentalization ↗

    D. Reuter, L. Marks, A. Carlucci, J. Ng, J. Petrie, J. Hausenloy, A. Karvonen, M. Baker · 2026 · ICML 2026 Workshop on Technical AI Governance Research

    Supports: workshop abstract proposes workload compartmentalization proofs to distinguish inference from frontier training; no protocol details assessed

    Locator: ICML 2026 TAIGR workshop abstract

    Version and catalogue details
  11. S-3565 / Tier A

    Shortcut-connected Expert Parallelism for Accelerating Mixture of Experts ↗

    W. Cai, J. Jiang, L. Qin, J. Cui, S. Kim, J. Huang · 2025 · ICML 2025, Proceedings of Machine Learning Research 267

    Supports: expert-parallel MoE inference involves all-to-all cross-device communication

    Locator: abstract

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review