M-0010 / On-chip & hardware

On-chip telemetry from timing, memory and performance counters

Uses timing, memory-residency and performance-counter signals measured on AI accelerators as evidence about which workloads they are running.

R2 DemonstratedSource reviewed 2026-09-25

01 / The mechanism and its boundary

What the technique establishes

On-chip telemetry uses signals measured on AI accelerators as evidence about their workloads, for example to tell training from inference or to check that model weights are held in a chip's own memory. Accelerators already track power, clock rates, memory use and operation counts, and challenge programs can time how quickly a chip completes set tasks. Studies on NVIDIA GPUs from the T4 to the B200 show that such signals can distinguish workloads. One classifier spotted training with 98.2% accuracy across its authors' corpus, but caught the hardest unseen disguised workloads only 43–87% of the time after five rounds of hardening. The biggest obstacle is trust in the data: an operator who controls the software stack can forge standard telemetry unless the read path is protected in hardware. The signals give statistical evidence, not proof. Profiling counters can also leak secrets, so NVIDIA disables them in full confidential-computing mode.

Threat model
Semi-trusted prover
Adversarial evaluation
Developer red-team
Hardware needed
Existing hardware features
Prover cooperation
Partial
Confidentiality
Partial
Category
On chip & hardware
Technical detail and cited results
  • Counter-based classification. Rahman and Tajdari sample nine always-available NVML counters at 1 Hz: GPU and memory utilization, memory used, power, temperature, SM and memory clocks, and PCIe TX/RX. They extract 166 features over 5–60 s windows, including memory slope and epoch periodicity from an FFT of power S-0037.
  • Memory-hard proof of work. In Monfared et al.'s challenge suite, these puzzles expose parallel effort and HBM use S-0033.
  • Verifiable delay functions. Based on sequential modular squaring, they expose sequential compute pressure S-0033.
  • GEMM puzzles. They target tensor-core throughput; the authors note that over 90% of LLM floating-point operations are GEMMs. Results can be checked with Freivalds' algorithm, subject to floating-point rounding discrepancies S-0033.
  • VRAM residency test. It runs bandwidth-bound Argon2id over challenge data. With 60 GB of challenge data on an H100, the response time for data held in HBM and for data in pinned host memory reached over PCIe differed by more than 350 ms S-0033.
  • Guaranteeable Memory. A guarantee chiplet beneath the HBM stacks would observe memory traffic directly. The author argues that the HBM standard makes it compatible with multiple leading accelerators S-0034.
  • Metering targets. Candidate targets for licensing include floating-point and integer arithmetic, memory, NVLink and PCIe transfer volume, energy and clock cycles S-0006.

Claims and scope

A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.

Readiness for a stated use

R2 Demonstrated

Assessed use: workload evidence from GPU counters and timing, assuming authentic measurements

medium confidence · current · assessed 2026-10-08 · rubric 1.1

This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.

Detailed results exist on current GPUs against a stated adversary, but the data can be forged in software and no one outside the developers has evaluated the methods.

  • R1 met: primitives, claims and threat models are published S-0033 S-0034 S-0037.
  • R2 met through reproducible published results. Rahman and Tajdari report end-to-end results on nine NVIDIA GPU models, including H100, H200 and B200, over five monitor-evader rounds against a stated adversary. They state that code for their workloads is in the project repository, but give no repository URL S-0037; the rubric does not require code for this route. Monfared et al. report measurements on T4 and H100 GPUs with language models of up to 72B parameters S-0033. No Implementation record realises this mechanism, and the most mature work is the NVML classifier.
  • R3 not met: no production-grade verification system built on these signals is available, and no source reports a party other than the developers relying on them for a verification decision. NVIDIA's Fleet Intelligence is a production telemetry service, but NVIDIA describes it as a tool for operators to monitor their own fleets S-3180. Trustworthy collection also needs hardware protections that Rahman and Tajdari list as prerequisites S-0037.
  • R4 not met: as of September 2026 no independent evaluation has been published.

Evidence needed for the next level

  • Use by a party other than the developers for a verification decision.

  • Telemetry read paths that the operator cannot forge, such as signed counters from a root of trust or a guarantee processor.

  • Calibrated false-positive and false-negative rates, with a detection-theoretic threshold framework.

  • Independent red-teaming, including custom-kernel and multi-node evasion.

Limitations, flaws, and blockers

These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.

critical / open / theoretical argument

Software-read telemetry can be forged by the operator

NVML-based classification assumes trustworthy telemetry. Without a tamper-resistant read path, an authenticated telemetry channel and secure boot of the monitoring software, an operator who controls the full software stack could forge counter values S-0037. Monfared et al. start from the same premise: current GPUs expose little trusted telemetry and can be modified or virtualized S-0033.

S-0037S-0033

significant / open / demonstrated attack

Adversarially disguised fine-tuning partly evades classification

Across 20 evasion strategy families in five monitor-evader rounds, the classifier's accuracy against the most challenging disguised workloads held out from training was 43–87%. White-box LoRA fine-tuning was the only evasion family detected less than half the time. The evaluation covered single nodes, PyTorch-level evasion and NVIDIA hardware S-0037.

S-0037

significant / open / theoretical argument

Timing challenges do not identify the individual chip

GEMM and VDF challenges can be answered by identical GPUs elsewhere, and floating-point fingerprints distinguish GPU models, not individual devices. GPU virtualization adds timing leakage that prevents attributing compute use S-0033.

S-0033

significant / open / theoretical argument

Counters leak information about protected workloads

Performance counters have been used as a side channel against TEEs, for example in CounterSEVeillance S-0014. NVIDIA disables performance counters in full confidential-computing mode, stating that they could provide an avenue for side-channel attacks S-1200. Richer counters for verification therefore pull against confidentiality.

S-0014S-1200

minor / open / open question

No quantified error rates or formal thresholds for timing primitives

Monfared et al. state that false-positive and false-negative rates are not quantified and leave hardware-specific formal thresholds to future work S-0033.

S-0033

What still blocks use or stronger assurance

  1. Shipping accelerators need a tamper-resistant, authenticated telemetry path.

    Dependency: Hardware-enabled guarantees (flexHEG) and guarantee processors

    S-0037S-0034
  2. NVIDIA's full confidential-computing mode disables the hardware performance counters its profiling tools use, so telemetry that needs them conflicts with it.

    S-1200S-0014
  3. Continuous challenge puzzles cost power and throughput on production workloads.

    S-0033
  4. Evaluation has not gone beyond single nodes, framework-level evasion and one vendor's hardware.

    S-0037

Connections in the research map

Depends on

Complementary techniques

Concepts used

Organizations and developers

The Consortium’s case files

Related editorial reviews use the Consortium’s own descriptive scores and review dates. Their scores are separate from the atlas readiness rubric.

CM-05 / AttestationFLOPs with a notary stampRead case file ↗

Sources and provenance

  1. S-0033 / Tier B

    Timing and Memory Telemetry on GPUs for AI Governance ↗

    S. K. Monfared, F. Ganji, D. E. Holcomb, S. Tajik · 2026 · arXiv

    Supports: four timing and memory primitives, threat model, T4/H100 results, residency-test conditions, overheads, limitations

    Locator: §3-§6, Figs. 5, 8, 10, 12, 15

    Version and catalogue details
  2. S-0034 / Tier B

    Guaranteeable Memory: An HBM-Based Chiplet for Verifiable AI Workloads ↗

    J. Petrie · 2025 · ICML 2025 Workshop on Technical AI Governance

    Supports: guarantee chiplet under HBM observing memory traffic; HBM-standard compatibility; independence from the accelerator die

    Locator: Abstract (read via ICML 2025 virtual site; OpenReview PDF not reachable)

    Version and catalogue details
  3. S-0037 / Tier B

    Detecting Hidden ML Training With Zero-Overhead Telemetry ↗

    R. Rahman, S. Tajdari · 2026 · ICML 2026 Workshop on Technical AI Governance Research

    Supports: NVML counter classifier, trust assumptions, GPU models, accuracy and evasion results, code statement, limitations

    Locator: Abstract; threat model; results; limitations

    Version and catalogue details
  4. S-0057 / Tier B

    Hardware-Enabled Governance Mechanisms: Developing Technical Solutions to Exempt Items Otherwise Classified Under Export Control Classification Numbers 3A090 and 4A090 ↗

    G. Kulp, D. Gonzales, E. Smith, L. Heim, P. Puri, M. J. D. Vermeer, Z. Winkelman · 2024 · RAND Corporation

    Supports: existing on-device counters and their use for metering

    Locator: p. 19

    Version and catalogue details
  5. S-0006 / Tier B

    Hardware-Enabled Mechanisms for Verifying Responsible AI Development ↗

    A. O'Gara, G. Kulp, W. Hodgkins, J. Petrie, V. Immler, A. Aysu, K. Basu, S. Bhasin, S. Picek, A. Srivastava · 2025 · arXiv

    Supports: candidate metering targets

    Locator: §2.2.2, §2.5.2, Table 1

    Version and catalogue details
  6. S-0014 / Tier C

    On TEEs for Privacy-Preserving Monitoring in AI Governance ↗

    Gloria Z · 2026 · MIRI Technical Governance Team

    Supports: counters as side channel; memory-residency and random challenges; completeness of workload declarations

    Version and catalogue details
  7. S-1200 / Tier B

    NVIDIA Secure AI with Blackwell and Hopper GPUs (White Paper) ↗

    NVIDIA · 2025 · NVIDIA documentation

    Supports: performance counters disabled in full CC mode, and NVIDIA's side-channel rationale

    Locator: p. 18

    Version and catalogue details
  8. S-0073 / Tier A

    Accurate and Convenient Energy Measurements for GPUs: A Detailed Study of NVIDIA GPU's Built-In Power Sensor ↗

    Z. Yang, K. Adamek, W. Armour · 2024 · SC24: International Conference for High Performance Computing, Networking, Storage and Analysis

    Supports: nvidia-smi power readings (via NVML) sample only 25% of runtime on A100 and H100; error about ±5% versus NVIDIA's claimed ±5 W

    Locator: Abstract; accuracy findings

    Version and catalogue details
  9. S-3180 / Tier B

    Introducing NVIDIA Fleet Intelligence for Real-Time GPU Fleet Visibility and Optimization ↗

    C. Shrauder, G. Frederick · 2026 · NVIDIA Technical Blog

    Supports: NVIDIA Fleet Intelligence: general availability, read-only open-source host agent, telemetry collected, signed attestation evidence (provider self-description)

    Locator: blog post

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review