M-0021 / Remote & side-channel sensing

Workload classification from telemetry and side channels

Telling whether chips are training, serving or doing non-AI work from GPU counters or power draw, signals that do not read weights or data.

R2 DemonstratedSource reviewed 2026-09-25

01 / The mechanism and its boundary

What the technique establishes

Workload classification infers what a chip is doing, such as training a model, serving it or running non-AI code. It uses signals that do not read weights or data: software counters such as power and utilization, or power draw measured outside the chip. Two 2026 preprints report high accuracy on their own test sets: 98.2% from software counters across nine NVIDIA GPU models, and 97% from an external probe on one H200. In the hardest cases, evasion strategies the classifier had not seen cut detection to 43–87% and 48–88% respectively; an added rule raised the second to at least 98%. The main obstacle is trust in the signal: software counters are read by software the operator controls, so they can be forged unless hardware protects the reading. External probes avoid this but have been tested on one GPU. Results cover single machines, and no independent evaluation exists.

Threat model
Adversarial prover
Adversarial evaluation
Developer red-team
Hardware needed
Existing hardware features
Prover cooperation
Partial
Confidentiality
Partial
Category
Remote & side channel sensing
Technical detail and cited results
  • NVML signals and corpus. Rahman and Tajdari sample nine NVML signals at about 1 Hz: GPU and memory utilization, memory used, power, temperature, SM and memory clocks, and PCIe transmit and receive bandwidth S-0037. Their corpus has 1,404 runs of 162 workloads, about 117 GPU-hours, on 9 NVIDIA GPU models across the Ampere, Ada, Hopper and Blackwell generations, with 1 to 8 GPUs per configuration; 421 of the runs are evasion attempts S-0037.
  • NVML classifier. A two-stage random forest, first ML versus non-ML and then training versus inference, classifies 30 s windows S-0037. Before adversarial hardening, five-fold cross-validation grouped by run (26,496 windows from 1,396 runs) gives 98.21% ± 0.84% accuracy, 97.39% precision, 98.90% recall and a 3.06% false-positive rate S-0037.
  • NVML evasion. The 20 evasion strategy families fall into six classes: utilization manipulation, memory reduction, workload mimicry, training disguised as inference, multi-GPU evasion and white-box attacks S-0037. In the fifth round the evader had white-box access to the previous round's classifier. After hardening, strategies seen in training were detected 100% of the time and held-out white-box strategies 43–87% of the time S-0037.
  • External power probe. Gargiulo and Kulp clamp a Rogowski current probe around the positive conductors of an NVIDIA H200 NVL's PCIe auxiliary power supply and digitize at a nominal 10 MHz in 5 s windows S-1412. Their dataset has 930 genuine and 680 adversarial recordings over 17 open model families and 25 non-AI workloads S-1412. On genuine recordings of model families unseen in training, they report 97% accuracy and a macro-averaged F1 of 0.955 across training, inference and non-AI classes S-1412.
  • Power-probe evasion. A detector hardened against evasion, with the tested strategy held out, caught training at least 99% of the time for three of four strategies. It caught diluted low-rank adaptation (LoRA) 48–88% of the time, rising to at least 98% with an added rule S-1412. The four strategies cost 28.1% to 69.3% of throughput S-1412. The authors state that these attacks are not a comprehensive evaluation against adversarial behaviour S-1412.

Claims and scope

A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.

Readiness for a stated use

R2 Demonstrated

Assessed use: telling training from inference and other work using genuine telemetry, including disguised workloads

medium confidence · current · assessed 2026-10-08 · rubric 1.1

This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.

Two preprints publish end-to-end results on current GPUs against stated, iterated adversaries; nobody independent has evaluated them.

  • R1 met: the claim, signals and threat model are publicly described S-0037 S-1412 S-0007.
  • R2 met: Rahman and Tajdari evaluate 9 current NVIDIA GPU models against 20 evasion strategy families over five rounds of monitor–evader iteration, with method and setup described in enough detail to repeat S-0037. They state that their workload source code is in the project repository, but the paper gives no link S-0037. Gargiulo and Kulp classify workloads on an H200 from externally measured power, test four evasion strategies and report releasing their recordings as a dataset S-1412.
  • R3 not met: no party other than the developers is reported to rely on these classifiers for a verification decision, and no production-grade system is available.
  • R4 not met: no independent adversarial evaluation has been published.

Confidence is medium: both results come from single machines, and the location of Rahman and Tajdari's code is not given.

Evidence needed for the next level

  • Reliance by a verifier other than the developers, or a production-grade, available system.

  • An independent red-team or peer-reviewed security analysis.

  • Results at multi-node cluster scale and across hardware vendors.

  • A tamper-resistant, authenticated telemetry path (see On-chip telemetry from timing, memory and performance counters), or physical sensing validated across devices.

Limitations, flaws, and blockers

These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.

significant / open / theoretical argument

Software-read counters can be forged by a privileged operator

Rahman and Tajdari state that without hardware-enabled guarantees, an adversary with software privilege can return forged counter values. Their results assume a tamper-resistant read path and an authenticated telemetry channel S-0037. Gargiulo and Kulp note that on-chip counters are read by software under the operator's control, so values could be reported without being measured, or replayed from an authorized workload while a hidden one runs S-1412. Against an operator who controls the full stack, forgery defeats classification from software counters, the signal source that On-chip telemetry from timing, memory and performance counters covers. It does not affect the external-probe route, which Gargiulo and Kulp argue can in principle be observed without the operator's cooperation S-1412.

S-0037S-1412

significant / open / demonstrated attack

Unseen evasion strategies still reduce detection

Rahman and Tajdari gave the evader white-box access to the previous round's classifier in their fifth round. After hardening against the other strategies of that round, their classifier detected the held-out white-box strategies only 43–87% of the time, against 100% for strategies it had been trained on S-0037. In Gargiulo and Kulp's power-trace study, a hardened detector that had not seen diluted LoRA caught it 48–88% of the time. An added rule raised detection to at least 98%, and the strategy cost the evader 33.8–69.3% of throughput S-1412.

S-0037S-1412

significant / open / theoretical argument

The training-inference boundary may blur

Ansari argues that as inference-time compute scaling and fine-tuning become more capable, the training-inference distinction may cease to be a reliable governance boundary S-0007.

S-0007

significant / open / open question

Limited scale and hardware coverage

Rahman and Tajdari's experiments are single-node, with up to 8 GPUs, and cover NVIDIA hardware only S-0037. Gargiulo and Kulp's corpus was collected on a single H200 NVL, and generalization across devices remains to be established S-1412.

S-0037S-1412

What still blocks use or stronger assurance

  1. Software telemetry is trustworthy only if on-chip counters are read over a path the operator cannot tamper with.

    Dependency: On-chip telemetry from timing, memory and performance counters

    S-0037S-1412
  2. No independent red-team or third-party reliance has been reported.

    S-0037S-1412
  3. Results do not yet cover multi-node clusters, other vendors or multi-tenant serving.

    S-0037S-1412

Connections in the research map

Depends on

Complementary techniques

Concepts used

Organizations and developers

The Consortium’s case files

Related editorial reviews use the Consortium’s own descriptive scores and review dates. Their scores are separate from the atlas readiness rubric.

CM-05 / AttestationFLOPs with a notary stampRead case file ↗

Sources and provenance

  1. S-0037 / Tier B

    Detecting Hidden ML Training With Zero-Overhead Telemetry ↗

    R. Rahman, S. Tajdari · 2026 · ICML 2026 Workshop on Technical AI Governance Research

    Supports: NVML-based classifier, corpus, cross-validated accuracy, evasion families and rounds, hardened detection of unseen strategies, threat model, trust assumption, code statement, limitations

    Locator: Abstract; §2.1, §2.2, §4.1-4.3, §5.1-5.2 and Table 5, §6.5; App. F

    Version and catalogue details
  2. S-1412 / Tier B

    Workload Identification with Physical Side Channels for AI Governance ↗

    S. Gargiulo, G. Kulp · 2026 · arXiv

    Supports: external power-probe classifier, accuracy on unseen model families, evasion strategies, hardened detection and costs, dataset release, NVML spoofing argument, limitations

    Locator: Abstract; §2-4; limitations

    Version and catalogue details
  3. S-0007 / Tier B

    Hardware-Level Governance of AI Compute: A Feasibility Taxonomy for Regulatory Compliance and Treaty Verification ↗

    S. Ansari · 2026 · arXiv

    Supports: workload-classification and power-monitoring feasibility; training-inference boundary

    Locator: §3.1 (M2, M4); §4.6

    Version and catalogue details
  4. S-0048 / Tier C

    Understanding Data Center Power Delivery ↗

    Amodo Design · 2026 · Amodo Design

    Supports: power delivery hierarchy filters signals; low-level monitoring harder to spoof

    Locator: whole note

    Version and catalogue details
  5. S-0041 / Tier A

    Single-Node Power Demand During AI Training: Measurements on an 8-GPU NVIDIA H100 System ↗

    I. Latif, A. C. Newkirk, M. R. Carbone, A. Munir, Y. Lin, J. Koomey, X. Yu, Z. Dong · 2025 · IEEE Access, vol. 13, pp. 61740–61747

    Supports: measured training power of an 8-GPU H100 node

    Locator: Abstract

    Version and catalogue details
  6. S-0042 / Tier B

    Input-Dependent Power Usage in GPUs ↗

    T. Gregersen, P. Patel, E. Choukse · 2024 · SC24-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis (Sustainable Supercomputing workshop), pp. 1872–1877

    Supports: input data changes GEMM power draw

    Locator: Abstract

    Version and catalogue details
  7. S-0040 / Tier A

    Detecting Covert Cryptomining Using HPC ↗

    A. Gangwal, S. G. Piazzetta, G. Lain, M. Conti · 2020 · Cryptology and Network Security – CANS 2020, LNCS 12579, pp. 344–364

    Supports: precedent: counter-based detection of covert cryptomining

    Locator: Abstract; evaluation

    Version and catalogue details
  8. S-0039 / Tier B

    Detecting Anomalies in Machine Learning Infrastructure via Hardware Telemetry ↗

    Z. Chen, S. Chien, P. Qian, N. Zilberman · 2025 · arXiv

    Supports: precedent: operator-accessible hardware signals for workload-agnostic anomaly detection

    Locator: Abstract; §3, §4.1

    Version and catalogue details
  9. S-0046 / Tier A

    DeepTheft: Stealing DNN Model Architectures through Power Side Channel ↗

    Y. Gao, H. Qiu, Z. Zhang, B. Wang, H. Ma, A. Abuadbba, M. Xue, A. Fu, S. Nepal · 2024 · 2024 IEEE Symposium on Security and Privacy

    Supports: power traces can leak model architecture

    Locator: Abstract

    Version and catalogue details
  10. S-0059 / Tier A

    Detecting Compute Structuring in AI Governance Is Likely Feasible ↗

    E. Seferis, T. Fist · 2026 · Proceedings of the AAAI Conference on Artificial Intelligence 40(44), pp. 37904–37912 (AAAI-26, Special Track on AI Alignment)

    Supports: compute-structuring detection: per-workload classification, aggregation of a customer's sequential or data-exchanging workloads against thresholds; analysis only; failure mode with very low data exchange

    Locator: threat models; Algorithms 1–2; limitations

    Version and catalogue details
  11. S-0073 / Tier A

    Accurate and Convenient Energy Measurements for GPUs: A Detailed Study of NVIDIA GPU's Built-In Power Sensor ↗

    Z. Yang, K. Adamek, W. Armour · 2024 · SC24: International Conference for High Performance Computing, Networking, Storage and Analysis

    Supports: nvidia-smi power readings (via NVML) sample only 25% of runtime on A100 and H100; error about ±5% versus NVIDIA's claimed ±5 W

    Locator: Abstract; accuracy findings

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review