01 / The mechanism and its boundary
What the technique establishes
On-chip telemetry uses signals measured on AI accelerators as evidence about their workloads, for example to tell training from inference or to check that model weights are held in a chip's own memory. Accelerators already track power, clock rates, memory use and operation counts, and challenge programs can time how quickly a chip completes set tasks. Studies on NVIDIA GPUs from the T4 to the B200 show that such signals can distinguish workloads. One classifier spotted training with 98.2% accuracy across its authors' corpus, but caught the hardest unseen disguised workloads only 43–87% of the time after five rounds of hardening. The biggest obstacle is trust in the data: an operator who controls the software stack can forge standard telemetry unless the read path is protected in hardware. The signals give statistical evidence, not proof. Profiling counters can also leak secrets, so NVIDIA disables them in full confidential-computing mode.
- Threat model
- Semi-trusted prover
- Adversarial evaluation
- Developer red-team
- Hardware needed
- Existing hardware features
- Prover cooperation
- Partial
- Confidentiality
- Partial
- Category
- On chip & hardware
Technical detail and cited results
- Counter-based classification. Rahman and Tajdari sample nine always-available NVML counters at 1 Hz: GPU and memory utilization, memory used, power, temperature, SM and memory clocks, and PCIe TX/RX. They extract 166 features over 5–60 s windows, including memory slope and epoch periodicity from an FFT of power S-0037.
- Memory-hard proof of work. In Monfared et al.'s challenge suite, these puzzles expose parallel effort and HBM use S-0033.
- Verifiable delay functions. Based on sequential modular squaring, they expose sequential compute pressure S-0033.
- GEMM puzzles. They target tensor-core throughput; the authors note that over 90% of LLM floating-point operations are GEMMs. Results can be checked with Freivalds' algorithm, subject to floating-point rounding discrepancies S-0033.
- VRAM residency test. It runs bandwidth-bound Argon2id over challenge data. With 60 GB of challenge data on an H100, the response time for data held in HBM and for data in pinned host memory reached over PCIe differed by more than 350 ms S-0033.
- Guaranteeable Memory. A guarantee chiplet beneath the HBM stacks would observe memory traffic directly. The author argues that the HBM standard makes it compatible with multiple leading accelerators S-0034.
- Metering targets. Candidate targets for licensing include floating-point and integer arithmetic, memory, NVLink and PCIe transfer volume, energy and clock cycles S-0006.
Claims and scope
A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.
This compute runs inference, not training
Classifying training versus other workloads from counters (S-0037) or timing (S-0033).
Declared hardware is idle or shut down
Utilization and residency signals indicate whether declared-idle hardware is busy (S-0033).
A training run stayed within declared limits
Counters for FLOP, memory and interconnect traffic are proposed as meters for compute accounting (S-0057, S-0006).
Readiness for a stated use
Assessed use: workload evidence from GPU counters and timing, assuming authentic measurements
medium confidence · current · assessed 2026-10-08 · rubric 1.1
This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.
Detailed results exist on current GPUs against a stated adversary, but the data can be forged in software and no one outside the developers has evaluated the methods.
- R1 met: primitives, claims and threat models are published S-0033 S-0034 S-0037.
- R2 met through reproducible published results. Rahman and Tajdari report end-to-end results on nine NVIDIA GPU models, including H100, H200 and B200, over five monitor-evader rounds against a stated adversary. They state that code for their workloads is in the project repository, but give no repository URL S-0037; the rubric does not require code for this route. Monfared et al. report measurements on T4 and H100 GPUs with language models of up to 72B parameters S-0033. No Implementation record realises this mechanism, and the most mature work is the NVML classifier.
- R3 not met: no production-grade verification system built on these signals is available, and no source reports a party other than the developers relying on them for a verification decision. NVIDIA's Fleet Intelligence is a production telemetry service, but NVIDIA describes it as a tool for operators to monitor their own fleets S-3180. Trustworthy collection also needs hardware protections that Rahman and Tajdari list as prerequisites S-0037.
- R4 not met: as of September 2026 no independent evaluation has been published.
Evidence needed for the next level
Use by a party other than the developers for a verification decision.
Telemetry read paths that the operator cannot forge, such as signed counters from a root of trust or a guarantee processor.
Calibrated false-positive and false-negative rates, with a detection-theoretic threshold framework.
Independent red-teaming, including custom-kernel and multi-node evasion.
Limitations, flaws, and blockers
These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.
critical / open / theoretical argument
Software-read telemetry can be forged by the operator
NVML-based classification assumes trustworthy telemetry. Without a tamper-resistant read path, an authenticated telemetry channel and secure boot of the monitoring software, an operator who controls the full software stack could forge counter values S-0037. Monfared et al. start from the same premise: current GPUs expose little trusted telemetry and can be modified or virtualized S-0033.
significant / open / demonstrated attack
Adversarially disguised fine-tuning partly evades classification
Across 20 evasion strategy families in five monitor-evader rounds, the classifier's accuracy against the most challenging disguised workloads held out from training was 43–87%. White-box LoRA fine-tuning was the only evasion family detected less than half the time. The evaluation covered single nodes, PyTorch-level evasion and NVIDIA hardware S-0037.
significant / open / theoretical argument
Timing challenges do not identify the individual chip
GEMM and VDF challenges can be answered by identical GPUs elsewhere, and floating-point fingerprints distinguish GPU models, not individual devices. GPU virtualization adds timing leakage that prevents attributing compute use S-0033.
significant / open / theoretical argument
Counters leak information about protected workloads
Performance counters have been used as a side channel against TEEs, for example in CounterSEVeillance S-0014. NVIDIA disables performance counters in full confidential-computing mode, stating that they could provide an avenue for side-channel attacks S-1200. Richer counters for verification therefore pull against confidentiality.
minor / open / open question
No quantified error rates or formal thresholds for timing primitives
Monfared et al. state that false-positive and false-negative rates are not quantified and leave hardware-specific formal thresholds to future work S-0033.
What still blocks use or stronger assurance
Shipping accelerators need a tamper-resistant, authenticated telemetry path.
Dependency: Hardware-enabled guarantees (flexHEG) and guarantee processors
S-0037S-0034- S-1200S-0014
NVIDIA's full confidential-computing mode disables the hardware performance counters its profiling tools use, so telemetry that needs them conflicts with it.
- S-0033
Continuous challenge puzzles cost power and throughput on production workloads.
- S-0037
Evaluation has not gone beyond single nodes, framework-level evasion and one vendor's hardware.
Connections in the research map
Depends on
- TEE remote attestation for AI workloads
A tamper-resistant read path, an authenticated channel and secure boot of the monitoring software are needed for counters to be trustworthy (S-0037).
Complementary techniques
Concepts used
Organizations and developers
The Consortium’s case files
Related editorial reviews use the Consortium’s own descriptive scores and review dates. Their scores are separate from the atlas readiness rubric.
CM-05 / AttestationFLOPs with a notary stampRead case file ↗Sources and provenance
- S-0033 / Tier B
Timing and Memory Telemetry on GPUs for AI Governance ↗
S. K. Monfared, F. Ganji, D. E. Holcomb, S. Tajik · 2026 · arXiv
Supports: four timing and memory primitives, threat model, T4/H100 results, residency-test conditions, overheads, limitations
Locator: §3-§6, Figs. 5, 8, 10, 12, 15
Version and catalogue details - S-0034 / Tier B
Guaranteeable Memory: An HBM-Based Chiplet for Verifiable AI Workloads ↗
J. Petrie · 2025 · ICML 2025 Workshop on Technical AI Governance
Supports: guarantee chiplet under HBM observing memory traffic; HBM-standard compatibility; independence from the accelerator die
Locator: Abstract (read via ICML 2025 virtual site; OpenReview PDF not reachable)
Version and catalogue details - S-0037 / Tier B
Detecting Hidden ML Training With Zero-Overhead Telemetry ↗
R. Rahman, S. Tajdari · 2026 · ICML 2026 Workshop on Technical AI Governance Research
Supports: NVML counter classifier, trust assumptions, GPU models, accuracy and evasion results, code statement, limitations
Locator: Abstract; threat model; results; limitations
Version and catalogue details - S-0057 / Tier B
Hardware-Enabled Governance Mechanisms: Developing Technical Solutions to Exempt Items Otherwise Classified Under Export Control Classification Numbers 3A090 and 4A090 ↗
G. Kulp, D. Gonzales, E. Smith, L. Heim, P. Puri, M. J. D. Vermeer, Z. Winkelman · 2024 · RAND Corporation
Supports: existing on-device counters and their use for metering
Locator: p. 19
Version and catalogue details - S-0006 / Tier B
Hardware-Enabled Mechanisms for Verifying Responsible AI Development ↗
A. O'Gara, G. Kulp, W. Hodgkins, J. Petrie, V. Immler, A. Aysu, K. Basu, S. Bhasin, S. Picek, A. Srivastava · 2025 · arXiv
Supports: candidate metering targets
Locator: §2.2.2, §2.5.2, Table 1
Version and catalogue details - S-0014 / Tier C
On TEEs for Privacy-Preserving Monitoring in AI Governance ↗
Gloria Z · 2026 · MIRI Technical Governance Team
Supports: counters as side channel; memory-residency and random challenges; completeness of workload declarations
Version and catalogue details - S-1200 / Tier B
NVIDIA Secure AI with Blackwell and Hopper GPUs (White Paper) ↗
NVIDIA · 2025 · NVIDIA documentation
Supports: performance counters disabled in full CC mode, and NVIDIA's side-channel rationale
Locator: p. 18
Version and catalogue details - S-0073 / Tier A
Accurate and Convenient Energy Measurements for GPUs: A Detailed Study of NVIDIA GPU's Built-In Power Sensor ↗
Z. Yang, K. Adamek, W. Armour · 2024 · SC24: International Conference for High Performance Computing, Networking, Storage and Analysis
Supports: nvidia-smi power readings (via NVML) sample only 25% of runtime on A100 and H100; error about ±5% versus NVIDIA's claimed ±5 W
Locator: Abstract; accuracy findings
Version and catalogue details - S-3180 / Tier B
Introducing NVIDIA Fleet Intelligence for Real-Time GPU Fleet Visibility and Optimization ↗
C. Shrauder, G. Frederick · 2026 · NVIDIA Technical Blog
Supports: NVIDIA Fleet Intelligence: general availability, read-only open-source host agent, telemetry collected, signed attestation evidence (provider self-description)
Locator: blog post
Version and catalogue details
- Source review date
- 2026-09-25
- Drafted by (source map)
- ai
- Review handles (source map)
- codex-review