01 / The mechanism and its boundary
What is being described
Interconnect bandwidth is the rate at which accelerators, servers or clusters can exchange data over the links between them; it is one of the measurable specifications of AI accelerators, alongside operations per second and memory capacity S-0053.
Large-scale training links thousands of accelerators with high-bandwidth interconnect, while efficient inference can run on dozens to low hundreds of closely connected accelerators S-0005. Between such pods, inference needs to pass only tokens, whereas training exchanges gradients or activations; Scher and Thiergart identify this gap as the target of bandwidth limits, as in bandwidth limits and compartmentalization S-0005. The gap holds only while each inference replica, including any split of the model or its experts across devices, stays within one pod. Mixture-of-experts inference that spreads experts across devices uses all-to-all communication between them S-3565. Inside a data centre, front-end links carry token-level inputs and outputs, while the back-end fabric between accelerators carries tensors and collective operations at much higher bandwidth, is latency-sensitive, and is harder to tap S-0018. US Executive Order 14110 defined reportable computing clusters partly by network connections faster than 100 Gbit/s S-0053. It was revoked in January 2025 S-0069. Sastry et al. note that the detectability of compute could be undermined if decentralized training, spread across many data centres or using lower-quality compute, becomes more viable S-0053.
Connections in the research map
Related research
Sources and provenance
- S-0053 / Tier B
Computing Power and the Governance of Artificial Intelligence ↗
G. Sastry, L. Heim, H. Belfield, M. Anderljung, M. Brundage, J. Hazell, C. O'Keefe, G. K. Hadfield, R. Ngo, K. Pilz, G. Gor, E. Bluemke, S. Shoker, J. Egan, R. F. Trager, S. Avin, A. Weller, Y. Bengio, D. Coyle · 2024 · arXiv
Supports: communication bandwidth as a chip specification alongside operations per second and memory; EO cluster definition using network connections over 100 Gbit/s; decentralized training risk
Locator: § on quantifiability and detectability; limitations
Version and catalogue details - S-0005 / Tier B
Mechanisms to Verify International Agreements About AI Development ↗
A. Scher, L. Thiergart · 2025 · arXiv
Supports: large-scale training links thousands of chips with high-bandwidth interconnect, efficient inference dozens to low hundreds; between pods inference needs tokens while training transfers gradients or activations; this gap is the target of bandwidth limits
Locator: Interconnect bandwidth limits
Version and catalogue details - S-3565 / Tier A
Shortcut-connected Expert Parallelism for Accelerating Mixture of Experts ↗
W. Cai, J. Jiang, L. Qin, J. Cui, S. Kim, J. Huang · 2025 · ICML 2025, Proceedings of Machine Learning Research 267
Supports: expert-parallel MoE inference involves all-to-all cross-device communication
Locator: abstract
Version and catalogue details - S-0018 / Tier B
A System Overview for Near-Term, Low-Trust AI Compute Verification ↗
N. Cankaya · 2026 · Machine Intelligence Research Institute
Supports: front-end token-level traffic vs high-bandwidth, latency-sensitive back-end fabric that is harder to tap
Locator: inference vs training
Version and catalogue details - S-0069 / Tier A
Executive Order 14148: Initial Rescissions of Harmful Executive Orders and Actions ↗
Executive Office of the President · 2025 · Federal Register, 90 FR 8237 (document 2025-01901, published 2025-01-28)
Supports: revocation of Executive Order 14110 in January 2025
Locator: §2(ggg)
Version and catalogue details
- Source review date
- 2026-09-25
- Drafted by (source map)
- ai
- Review handles (source map)
- codex-review