K-0025

Inference and training workloads

Training updates a model's weights from data. Inference uses fixed weights to produce outputs. Their different resource use underpins several verification methods.

Source reviewed 2026-09-25

01 / The mechanism and its boundary

What is being described

Training is the workload that updates a model's weights step by step from batches of data, and inference is the workload that runs a model with fixed weights on inputs to produce outputs such as tokens S-0029 S-0018.

RAND's verification framework treats declared training and declared inference as distinct uses of compute, each to be verified S-0002. Sastry and colleagues note that most AI compute is used for inference, although a single training run needs far more compute than a single inference, and individual copies of a model can run on relatively little compute S-0053. Verification designs use the differences in resource use:

  • Large-scale training links thousands of accelerators and exchanges gradients or activations between groups of them, while inference between pods passes only tokens S-0005, provided each inference replica, including any split of the model across devices, stays within one pod S-3565. Bandwidth limits rely on this gap.
  • Training and inference often differ in accelerator utilization and power draw S-0005, which workload classification from telemetry and side channels uses.
  • One classifier using GPU telemetry reports 98.2% binary accuracy at identifying training across its corpus of nine GPU models, falling to 43–87% on the most challenging disguised workloads held out from its training S-0037, as in on-chip telemetry.

Reinforcement learning blurs the line, because its rollouts are inference. JoshC analyzes a covert strategy that generates rollouts on declared inference servers while updating the model on hidden compute S-3566. Shavit notes that there is no straightforward way to tell whether an accelerator is running training or an unrelated workload S-0029.

Connections in the research map

Related research

Sources and provenance

  1. S-0029 / Tier B

    What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring ↗

    Y. Shavit · 2023 · arXiv

    Supports: training steps update weights from data batches; no straightforward way to tell whether an ML chip is running training or an unrelated job

    Locator: §5.1; §4

    Version and catalogue details
  2. S-0018 / Tier B

    A System Overview for Near-Term, Low-Trust AI Compute Verification ↗

    N. Cankaya · 2026 · Machine Intelligence Research Institute

    Supports: inference yields token-level input-output traffic on front-end links; training uses the back-end fabric

    Locator: inference vs training

    Version and catalogue details
  3. S-0002 / Tier B

    Verifying International Agreements on AI: Six Layers of Verification for Rules on Large-Scale AI Development and Deployment ↗

    M. Baker, G. Kulp, O. Marks, M. Brundage, L. Heim · 2025 · RAND Corporation

    Supports: declared training (1.A.1) and inference (1.A.2) as distinct declared uses

    Locator: §3.2

    Version and catalogue details
  4. S-0053 / Tier B

    Computing Power and the Governance of Artificial Intelligence ↗

    G. Sastry, L. Heim, H. Belfield, M. Anderljung, M. Brundage, J. Hazell, C. O'Keefe, G. K. Hadfield, R. Ngo, K. Pilz, G. Gor, E. Bluemke, S. Shoker, J. Egan, R. F. Trager, S. Avin, A. Weller, Y. Bengio, D. Coyle · 2024 · arXiv

    Supports: majority of AI compute used for inference; single training run needs far more compute than a single inference; copies of a model run on little compute

    Locator: training vs inference discussion

    Version and catalogue details
  5. S-0005 / Tier B

    Mechanisms to Verify International Agreements About AI Development ↗

    A. Scher, L. Thiergart · 2025 · arXiv

    Supports: large-scale training links thousands of chips and exchanges gradients; efficient inference on dozens to low hundreds of chips passes only tokens between pods; utilization and power often differ

    Locator: Interconnect bandwidth limits; workload classification with high-level chip measures

    Version and catalogue details
  6. S-3565 / Tier A

    Shortcut-connected Expert Parallelism for Accelerating Mixture of Experts ↗

    W. Cai, J. Jiang, L. Qin, J. Cui, S. Kim, J. Huang · 2025 · ICML 2025, Proceedings of Machine Learning Research 267

    Supports: expert-parallel MoE inference involves all-to-all cross-device communication

    Locator: abstract

    Version and catalogue details
  7. S-3566 / Tier C

    Can governments quickly and cheaply slow AI training? ↗

    joshc · 2026 · AI Alignment Forum

    Supports: covert reinforcement-learning rollouts on declared inference servers with updates on hidden compute

    Locator: §3.4; §4

    Version and catalogue details
  8. S-0037 / Tier B

    Detecting Hidden ML Training With Zero-Overhead Telemetry ↗

    R. Rahman, S. Tajdari · 2026 · ICML 2026 Workshop on Technical AI Governance Research

    Supports: NVML-telemetry classifier: 98.2% binary accuracy at identifying training across its corpus (9 GPU models); 43–87% against the white-box disguised workloads held out from training, after hardening

    Locator: abstract

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review