01 / The mechanism and its boundary
What is being described
The claim is that a declared cluster is used only to run existing models to produce outputs, not to train new or more capable models. Several agreement proposals would let existing AI models keep serving users while restricting further training. The claim combines a positive part (the declared inference is happening) with a negative part (nothing else, especially training, is). It matters because it could let most AI compute stay in productive use under a training restriction. It is hard because the same chips can do both, workloads can be disguised, and the negative part requires accounting for all of a facility's activity. Proposed approaches include capturing and recomputing the traffic that enters and leaves a facility, limiting bandwidth between groups of chips so that large training cannot be coordinated, and classifying workloads from telemetry. Each rests on open assumptions about numerical nondeterminism, hidden capacity and side channels.
State of verification
Editorial synthesis from the AI Verification Tech Map.
Declared inference can be checked with components that are demonstrated (R2) or in production (R3), but the absence of training cannot yet be verified. That half rests on full-stack designs (AI 2040 inference-only verification stack, Low-trust AI compute verification system overview, SASH confidential network logger) that are proposed (R1), and a team building the components reports nothing past a proof-of-principle prototype S-1512.
For example, a verifier may want to know that a cluster declared for inference is not training a new model.
- Network taps (R1) record the cluster's front-end traffic with its users, and sampled recomputation re-runs sampled requests on the declared model. DiFR (R2) tolerates numerical noise in these re-runs on open-weight models of 8 to 30 billion parameters S-0016.
- SASH's network logger (R1) is a public prototype of both. It passes every request through a logger and re-runs it on a separate cluster, with a 270-million-parameter model and no stated adversary S-1319.
- This covers the positive half at most. It shows that sampled outputs match the declared model. A tap on the cluster's external links does not stop covert workloads. It aims only to stop their results leaving over those links S-1300.
- The negative half needs the rest of the cluster accounted for. Training traffic runs on back-end fabric that is harder to tap S-0018. Bandwidth limits (R2 for software monitoring) between groups of chips target that fabric, since distributed training must exchange gradients S-0005. Attestable proposes proofs of useful work (R1) to keep declared hardware busy with approved or protocol-defined work, leaving little spare capacity S-1102.
- Known routes remain. Reinforcement-learning rollouts are inference, so declared servers could generate them while hidden compute updates the model S-3566. In one scenario, JoshC estimates that more than 95% of computation must be accounted for to constrain that strategy S-3566. Workload classification (R2) is a lighter alternative. It detects training with 98.2% accuracy on its own corpus, but 43–87% on the most challenging disguised workloads held out from its training, and current GPUs lack the protections its telemetry needs to be trusted S-0037 S-0033.
Connections in the research map
Concepts used
Techniques addressing this claim
Sources and provenance
- S-0063 / Tier B
An International Agreement to Prevent the Premature Creation of Artificial Superintelligence ↗
A. Scher, D. Abecassis, P. Barnett, B. Abeyta · 2025 · Machine Intelligence Research Institute
Supports: restricting the scale of training; chip use verification distinguishing inference on existing systems from training
Locator: abstract; Article VII (as summarised)
Version and catalogue details - S-0067 / Tier C
Verification Plan ↗
R. Dean · 2026 · AI 2040
Supports: data centres converted to inference-only operation; network taps, recomputation and reproducible packets
Locator: phases; verification mechanisms
Version and catalogue details - S-0053 / Tier B
Computing Power and the Governance of Artificial Intelligence ↗
G. Sastry, L. Heim, H. Belfield, M. Anderljung, M. Brundage, J. Hazell, C. O'Keefe, G. K. Hadfield, R. Ngo, K. Pilz, G. Gor, E. Bluemke, S. Shoker, J. Egan, R. F. Trager, S. Avin, A. Weller, Y. Bengio, D. Coyle · 2024 · arXiv
Supports: majority of AI compute used for inference; decentralised training could undermine detectability
Locator: training vs inference; limitations
Version and catalogue details - S-0002 / Tier B
Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment ↗
M. Baker, G. Kulp, O. Marks, M. Brundage, L. Heim · 2025 · RAND Corporation
Supports: declared inference (1.A.2) as a distinct subgoal; deterministic replication of inference as an R&D problem
Locator: §3.2; Appendix A.9
Version and catalogue details - S-0018 / Tier B
A System Overview for Near-Term, Low-Trust AI Compute Verification ↗
N. Cankaya · 2026 · Machine Intelligence Research Institute
Supports: distinguishing inference from training; token-level front-end evidence; back-end harder to tap; egress explainable by ingress as open question; memory wiping; side channels
Locator: verification goals; inference vs training; open problems
Version and catalogue details - S-0005 / Tier B
Mechanisms to Verify International Agreements About AI Development ↗
A. Scher, L. Thiergart · 2025 · arXiv
Supports: inference-specialised chips repurposable for training; pods with limited external bandwidth
Locator: Verifying that known compute is not being used for a large training run
Version and catalogue details - S-0029 / Tier B
What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring ↗
Y. Shavit · 2023 · arXiv
Supports: no straightforward way to tell whether a chip is running training or another job
Locator: open problems
Version and catalogue details - S-0037 / Tier B
Detecting Hidden ML Training With Zero-Overhead Telemetry ↗
R. Rahman, S. Tajdari · 2026 · ICML 2026 Workshop on Technical AI Governance Research
Supports: telemetry classifier accuracy overall and on adversarially disguised workloads; required telemetry protections
Locator: abstract; §5.2; deployment requirements
Version and catalogue details - S-0001 / Tier A
Open Problems in Technical AI Governance ↗
A. Reuel, B. Bucknall, S. Casper, T. Fist, L. Soder, O. Aarne, L. Hammond, L. Ibrahim, A. Chan, P. Wills, M. Anderljung, B. Garfinkel, L. Heim, A. Trask, G. Mukobi, R. Schaeffer, M. Baker, S. Hooker, I. Solaiman, A. S. Luccioni, N. Rajkumar, N. Moës, J. Ladish, D. Bau, P.-A. Bricman, N. Guha, J. Newman, Y. Bengio, T. South, A. Pentland, S. Koyejo, M. J. Kochenderfer, R. Trager · 2025 · Transactions on Machine Learning Research
Supports: workload classification; adversarial customers may obfuscate by adding noise
Locator: §3.2.2 / §5.2.2 open problems
Version and catalogue details - S-0031 / Tier C
The Fundamentals and Feasibility of Secure Network Taps for Verifying AI Datacenter Use ↗
N. Cankaya · 2026 · The Datacenter Lie Detector
Supports: front-end tapping most viable; back-end requires sampling
Locator: frontend vs backend
Version and catalogue details - S-0016 / Tier B
DiFR: Inference Verification Despite Nondeterminism ↗
A. Karvonen, D. Reuter, R. Rinberg, L. Marks, A. Garriga-Alonso, K. Warr · 2025 · ICML 2026 Workshop on Technical AI Governance Research
Supports: recomputation despite benign numerical noise on 8–30B open-weight models
Locator: abstract; §5
Version and catalogue details - S-0020 / Tier B
Bit-Exact AI Inference Verification Without Performance Tradeoffs ↗
N. Cankaya · 2026 · ICML 2026 Workshop on Technical AI Governance Research
Supports: bit-exact inference verification across GPU variants
Locator: abstract
Version and catalogue details - S-1512 / Tier C
Verifying international AI deals: Plan A, the state-of-play, and what you can do to help ↗
T. Milton, S. Reynolds, C. Jacobi, J. Foster · 2026 · Amodo (Substack)
Supports: no verification component past a proof-of-principle prototype
Locator: introduction
Version and catalogue details - S-0033 / Tier B
Timing and Memory Telemetry on GPUs for AI Governance ↗
S. K. Monfared, F. Ganji, D. E. Holcomb, S. Tajik · 2026 · arXiv
Supports: current GPUs expose limited trusted telemetry
Locator: abstract
Version and catalogue details - S-1102 / Tier C
Pacing AI Requires Proof ↗
Attestable · 2026 · Attestable blog
Supports: spare capacity could run an unauthorised training job; approved inference plus protocol-defined work fills a required work budget (provider proposal)
Locator: blog post
Version and catalogue details - S-3566 / Tier C
Can governments quickly and cheaply slow AI training? ↗
joshc · 2026 · AI Alignment Forum
Supports: public analysis of covert reinforcement-learning rollouts on declared inference servers and updates on hidden compute; share of computation that must be accounted for (scenario estimate)
Locator: §2.5; §3.4; §4
Version and catalogue details - S-1319 / Tier B
inference-verification: Inference Verification Prototype ↗
Singapore AI Safety Hub (SASH) · 2026 · GitHub
Supports: SASH prototype re-runs every request through a logger on a separate cluster, with Gemma 3 270M and a demo switch as the only adversary
Locator: README.md; implementation
Version and catalogue details - S-1300 / Tier B
Fingerprinting All AI Cluster I/O Without Mutually Trusted Processors ↗
N. Cankaya, J. Kryś, J. Ng, L. Marks, F. Krückel · 2026 · arXiv
Supports: a tap on external links does not prevent covert workloads, only the exfiltration of their results
Version and catalogue details
- Source review date
- 2026-09-25
- Drafted by (source map)
- ai
- Review handles (source map)
- codex-review