K-0017

Compartmentalization

Dividing a facility's accelerators into groups with restricted communication between them, so that combining groups for large training runs becomes slow or impractical.

Source reviewed 2026-09-25

01 / The mechanism and its boundary

What is being described

Compartmentalization divides a facility's accelerators into groups and restricts communication between the groups, so that groups are hard to combine into a larger workload than is allowed S-0005 S-0057.

Scher and Thiergart describe pods of well-connected chips. Between pods, inference needs to pass only tokens, whereas other forms of parallelism transfer activations or gradients S-0005. That holds only while each inference replica, including any split of the model or its experts across devices, stays within one pod. Mixture-of-experts inference that spreads experts across devices uses all-to-all communication between them S-3565. Efficient inference fits within dozens to low hundreds of closely connected accelerators, while large-scale training links thousands S-0005. RAND's "fixed set" design likewise restricts networking so that small, fixed sets of GPUs cannot be combined into large clusters S-0057, and Sastry and colleagues list physical limits on chip-to-chip networking as a way to enforce compute caps S-0053. These ideas underlie bandwidth limits and compartmentalization.

Compartments can also separate trust domains. One low-trust design air-gaps its evaluation environments and uses optical splitters and data diodes, simple components that can be inspected for tampering, to enforce one-way data movement S-0018. The boundaries can be checked by observing traffic between accelerators with network taps S-0002, while physical channels that could bypass monitored links are the target of side-channel suppression S-0038.

Connections in the research map

Related research

Sources and provenance

  1. S-0005 / Tier B

    Mechanisms to Verify International Agreements About AI Development ↗

    A. Scher, L. Thiergart · 2025 · arXiv

    Supports: pods of well-connected chips; between pods, inference needs only tokens while other forms of parallelism transfer activations or gradients; efficient inference on dozens to low hundreds of chips versus thousands for large training

    Locator: Interconnect bandwidth limits

    Version and catalogue details
  2. S-3565 / Tier A

    Shortcut-connected Expert Parallelism for Accelerating Mixture of Experts ↗

    W. Cai, J. Jiang, L. Qin, J. Cui, S. Kim, J. Huang · 2025 · ICML 2025, Proceedings of Machine Learning Research 267

    Supports: expert-parallel MoE inference involves all-to-all cross-device communication

    Locator: abstract

    Version and catalogue details
  3. S-0057 / Tier B

    Hardware-Enabled Governance Mechanisms: Developing Technical Solutions to Exempt Items Otherwise Classified Under Export Control Classification Numbers 3A090 and 4A090 ↗

    G. Kulp, D. Gonzales, E. Smith, L. Heim, P. Puri, M. J. D. Vermeer, Z. Winkelman · 2024 · RAND Corporation

    Supports: fixed-set HEM restricting networking of small, fixed sets of GPUs

    Locator: p. viii

    Version and catalogue details
  4. S-0053 / Tier B

    Computing Power and the Governance of Artificial Intelligence ↗

    G. Sastry, L. Heim, H. Belfield, M. Anderljung, M. Brundage, J. Hazell, C. O'Keefe, G. K. Hadfield, R. Ngo, K. Pilz, G. Gor, E. Bluemke, S. Shoker, J. Egan, R. F. Trager, S. Avin, A. Weller, Y. Bengio, D. Coyle · 2024 · arXiv

    Supports: compute caps enforced via physical limits on chip-to-chip networking

    Locator: enforcement mechanisms

    Version and catalogue details
  5. S-0018 / Tier B

    A System Overview for Near-Term, Low-Trust AI Compute Verification ↗

    N. Cankaya · 2026 · Machine Intelligence Research Institute

    Supports: air-gapped evaluation environments; optical splitters and data diodes as inspectable components enforcing one-way data movement

    Locator: system architecture

    Version and catalogue details
  6. S-0002 / Tier B

    Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment ↗

    M. Baker, G. Kulp, O. Marks, M. Brundage, L. Heim · 2025 · RAND Corporation

    Supports: network taps observing data exchanged between chips

    Locator: §4.2

    Version and catalogue details
  7. S-0038 / Tier C

    Suppressing Side Channels in an Untrusted Data Center via Retrofitted Defenses ↗

    N. Cankaya · 2026 · MIRI Technical Governance Team

    Supports: physical channels could bypass network monitoring

    Locator: side channels of concern

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review