C-0004

This compute runs inference, not training

A declared cluster is used only to run existing models to produce outputs, and not to train new or more capable models.

Source reviewed 2026-09-25

01 / The mechanism and its boundary

What is being described

The claim is that a declared cluster is used only to run existing models to produce outputs, not to train new or more capable models. Several agreement proposals would let existing AI models keep serving users while restricting further training. The claim combines a positive part (the declared inference is happening) with a negative part (nothing else, especially training, is). It matters because it could let most AI compute stay in productive use under a training restriction. It is hard because the same chips can do both, workloads can be disguised, and the negative part requires accounting for all of a facility's activity. Proposed approaches include capturing and recomputing the traffic that enters and leaves a facility, limiting bandwidth between groups of chips so that large training cannot be coordinated, and classifying workloads from telemetry. Each rests on open assumptions about numerical nondeterminism, hidden capacity and side channels.

State of verification

Editorial synthesis from the AI Verification Tech Map.

Declared inference can be checked with components that are demonstrated (R2) or in production (R3), but the absence of training cannot yet be verified. That half rests on full-stack designs (AI 2040 inference-only verification stack, Low-trust AI compute verification system overview, SASH confidential network logger) that are proposed (R1), and a team building the components reports nothing past a proof-of-principle prototype S-1512.

For example, a verifier may want to know that a cluster declared for inference is not training a new model.

  • Network taps (R1) record the cluster's front-end traffic with its users, and sampled recomputation re-runs sampled requests on the declared model. DiFR (R2) tolerates numerical noise in these re-runs on open-weight models of 8 to 30 billion parameters S-0016.
  • SASH's network logger (R1) is a public prototype of both. It passes every request through a logger and re-runs it on a separate cluster, with a 270-million-parameter model and no stated adversary S-1319.
  • This covers the positive half at most. It shows that sampled outputs match the declared model. A tap on the cluster's external links does not stop covert workloads. It aims only to stop their results leaving over those links S-1300.
  • The negative half needs the rest of the cluster accounted for. Training traffic runs on back-end fabric that is harder to tap S-0018. Bandwidth limits (R2 for software monitoring) between groups of chips target that fabric, since distributed training must exchange gradients S-0005. Attestable proposes proofs of useful work (R1) to keep declared hardware busy with approved or protocol-defined work, leaving little spare capacity S-1102.
  • Known routes remain. Reinforcement-learning rollouts are inference, so declared servers could generate them while hidden compute updates the model S-3566. In one scenario, JoshC estimates that more than 95% of computation must be accounted for to constrain that strategy S-3566. Workload classification (R2) is a lighter alternative. It detects training with 98.2% accuracy on its own corpus, but 43–87% on the most challenging disguised workloads held out from its training, and current GPUs lack the protections its telemetry needs to be trusted S-0037 S-0033.

Connections in the research map

Concepts used

Techniques addressing this claim

Sources and provenance

  1. S-0063 / Tier B

    An International Agreement to Prevent the Premature Creation of Artificial Superintelligence ↗

    A. Scher, D. Abecassis, P. Barnett, B. Abeyta · 2025 · Machine Intelligence Research Institute

    Supports: restricting the scale of training; chip use verification distinguishing inference on existing systems from training

    Locator: abstract; Article VII (as summarised)

    Version and catalogue details
  2. S-0067 / Tier C

    Verification Plan ↗

    R. Dean · 2026 · AI 2040

    Supports: data centres converted to inference-only operation; network taps, recomputation and reproducible packets

    Locator: phases; verification mechanisms

    Version and catalogue details
  3. S-0053 / Tier B

    Computing Power and the Governance of Artificial Intelligence ↗

    G. Sastry, L. Heim, H. Belfield, M. Anderljung, M. Brundage, J. Hazell, C. O'Keefe, G. K. Hadfield, R. Ngo, K. Pilz, G. Gor, E. Bluemke, S. Shoker, J. Egan, R. F. Trager, S. Avin, A. Weller, Y. Bengio, D. Coyle · 2024 · arXiv

    Supports: majority of AI compute used for inference; decentralised training could undermine detectability

    Locator: training vs inference; limitations

    Version and catalogue details
  4. S-0002 / Tier B

    Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment ↗

    M. Baker, G. Kulp, O. Marks, M. Brundage, L. Heim · 2025 · RAND Corporation

    Supports: declared inference (1.A.2) as a distinct subgoal; deterministic replication of inference as an R&D problem

    Locator: §3.2; Appendix A.9

    Version and catalogue details
  5. S-0018 / Tier B

    A System Overview for Near-Term, Low-Trust AI Compute Verification ↗

    N. Cankaya · 2026 · Machine Intelligence Research Institute

    Supports: distinguishing inference from training; token-level front-end evidence; back-end harder to tap; egress explainable by ingress as open question; memory wiping; side channels

    Locator: verification goals; inference vs training; open problems

    Version and catalogue details
  6. S-0005 / Tier B

    Mechanisms to Verify International Agreements About AI Development ↗

    A. Scher, L. Thiergart · 2025 · arXiv

    Supports: inference-specialised chips repurposable for training; pods with limited external bandwidth

    Locator: Verifying that known compute is not being used for a large training run

    Version and catalogue details
  7. S-0029 / Tier B

    What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring ↗

    Y. Shavit · 2023 · arXiv

    Supports: no straightforward way to tell whether a chip is running training or another job

    Locator: open problems

    Version and catalogue details
  8. S-0037 / Tier B

    Detecting Hidden ML Training With Zero-Overhead Telemetry ↗

    R. Rahman, S. Tajdari · 2026 · ICML 2026 Workshop on Technical AI Governance Research

    Supports: telemetry classifier accuracy overall and on adversarially disguised workloads; required telemetry protections

    Locator: abstract; §5.2; deployment requirements

    Version and catalogue details
  9. S-0001 / Tier A

    Open Problems in Technical AI Governance ↗

    A. Reuel, B. Bucknall, S. Casper, T. Fist, L. Soder, O. Aarne, L. Hammond, L. Ibrahim, A. Chan, P. Wills, M. Anderljung, B. Garfinkel, L. Heim, A. Trask, G. Mukobi, R. Schaeffer, M. Baker, S. Hooker, I. Solaiman, A. S. Luccioni, N. Rajkumar, N. Moës, J. Ladish, D. Bau, P.-A. Bricman, N. Guha, J. Newman, Y. Bengio, T. South, A. Pentland, S. Koyejo, M. J. Kochenderfer, R. Trager · 2025 · Transactions on Machine Learning Research

    Supports: workload classification; adversarial customers may obfuscate by adding noise

    Locator: §3.2.2 / §5.2.2 open problems

    Version and catalogue details
  10. S-0031 / Tier C

    The Fundamentals and Feasibility of Secure Network Taps for Verifying AI Datacenter Use ↗

    N. Cankaya · 2026 · The Datacenter Lie Detector

    Supports: front-end tapping most viable; back-end requires sampling

    Locator: frontend vs backend

    Version and catalogue details
  11. S-0016 / Tier B

    DiFR: Inference Verification Despite Nondeterminism ↗

    A. Karvonen, D. Reuter, R. Rinberg, L. Marks, A. Garriga-Alonso, K. Warr · 2025 · ICML 2026 Workshop on Technical AI Governance Research

    Supports: recomputation despite benign numerical noise on 8–30B open-weight models

    Locator: abstract; §5

    Version and catalogue details
  12. S-0020 / Tier B

    Bit-Exact AI Inference Verification Without Performance Tradeoffs ↗

    N. Cankaya · 2026 · ICML 2026 Workshop on Technical AI Governance Research

    Supports: bit-exact inference verification across GPU variants

    Locator: abstract

    Version and catalogue details
  13. S-1512 / Tier C

    Verifying international AI deals: Plan A, the state-of-play, and what you can do to help ↗

    T. Milton, S. Reynolds, C. Jacobi, J. Foster · 2026 · Amodo (Substack)

    Supports: no verification component past a proof-of-principle prototype

    Locator: introduction

    Version and catalogue details
  14. S-0033 / Tier B

    Timing and Memory Telemetry on GPUs for AI Governance ↗

    S. K. Monfared, F. Ganji, D. E. Holcomb, S. Tajik · 2026 · arXiv

    Supports: current GPUs expose limited trusted telemetry

    Locator: abstract

    Version and catalogue details
  15. S-1102 / Tier C

    Pacing AI Requires Proof ↗

    Attestable · 2026 · Attestable blog

    Supports: spare capacity could run an unauthorised training job; approved inference plus protocol-defined work fills a required work budget (provider proposal)

    Locator: blog post

    Version and catalogue details
  16. S-3566 / Tier C

    Can governments quickly and cheaply slow AI training? ↗

    joshc · 2026 · AI Alignment Forum

    Supports: public analysis of covert reinforcement-learning rollouts on declared inference servers and updates on hidden compute; share of computation that must be accounted for (scenario estimate)

    Locator: §2.5; §3.4; §4

    Version and catalogue details
  17. S-1319 / Tier B

    inference-verification: Inference Verification Prototype ↗

    Singapore AI Safety Hub (SASH) · 2026 · GitHub

    Supports: SASH prototype re-runs every request through a logger on a separate cluster, with Gemma 3 270M and a demo switch as the only adversary

    Locator: README.md; implementation

    Version and catalogue details
  18. S-1300 / Tier B

    Fingerprinting All AI Cluster I/O Without Mutually Trusted Processors ↗

    N. Cankaya, J. Kryś, J. Ng, L. Marks, F. Krückel · 2026 · arXiv

    Supports: a tap on external links does not prevent covert workloads, only the exfiltration of their results

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review