M-0011 / On-chip & hardware

Hardware performance throttling and licensing

On-chip mechanisms that cut an AI accelerator's performance when a license expires or a trusted trigger fires, bounding what the hardware can do.

R1 ProposedSource reviewed 2026-09-25

01 / The mechanism and its boundary

What the technique establishes

Throttling mechanisms would let the hardware itself enforce limits on AI computation. One proposal is offline licensing: a chip runs at full speed only while it holds a valid license for a set amount of work. Another is a set of microarchitectural throttles that, when triggered, shrink L2 cache capacity, add cache latency, cap cache bandwidth or limit shared-memory access. A 2026 simulation of an NVIDIA A100 found that such throttles could cut the performance of large-language-model kernels by up to 80% at one-eighth of resource availability, at a cost of under about 10,000 flip-flops each. The results come from simulation, not real chips. For verification, throttling matters only if a verifier can confirm the limit is in place and cannot be bypassed. That depends on attestation, tamper resistance and a secure trigger or licensing path. RAND judges that anti-tamper protection would not be insurmountable for a determined, well-resourced adversary.

Threat model
Adversarial prover
Adversarial evaluation
Published analysis
Hardware needed
New chip design
Prover cooperation
Partial
Confidentiality
Preserving
Category
On chip & hardware
Technical detail and cited results
  • Knobs evaluated by Ma et al. L2 capacity through way masking, L2 latency through configurable request buffering, L2 bandwidth through credit-based rate limiting, and shared-memory port access through bank arbitration. None is exposed to a software-visible interface S-0036.
  • Setup. The simulator was AccelSim, configured as an NVIDIA A100, with utilization metrics validated on a real GPU with Nsight Compute. The workloads were CUTLASS StreamK GEMM kernels, for prefill (M=4096) and decode (M=128), and FlashAttention, with shapes from DeepSeek-V3, Llama-3-70B and Mixtral-8x7B, plus six non-LLM CUDA sample workloads S-0036.
  • Results. At one-eighth of resource availability, performance fell by up to 80% (L2 latency for decode; L2 associativity for prefill). After a throttle was applied, performance settled within about 5K cycles for shared-memory ports, 5–7K cycles for L2 response rate and about 80K cycles for L2 associativity. Each mechanism needs fewer than about 10K flip-flops S-0036.
  • Selectivity. Memory-side knobs affected LLM kernels selectively. Disabling compute cores degraded nearly every workload together S-0036.
  • Offline licensing (RAND). A secure message authenticated on the GPU might authorize, for example, 10^18 arithmetic operations, after which the GPU would fall back to 1% of full performance S-0057.
  • Embedded off-switch (Petrie). Each security block uses public-key signatures, with nonces against replay, to check for recent authorization. Petrie estimates roughly 40,000 transistors per block, and about 0.5% of the die for 10,000 blocks S-3160. The proof-of-concept README reports signature checks of about 5 million cycles for ECDSA and about 300,000 cycles for HSS/LMS S-3161.

Claims and scope

A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.

Readiness for a stated use

R1 Proposed

Assessed use: performance limits a verifier can rely on, against an operator trying to bypass them

medium confidence · current · assessed 2026-09-25 · rubric 1.1

This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.

The designs are published, but the only quantitative results are kernel-level simulations and the only public code is a simulated proof of concept.

  • R1 met: RAND and O'Gara et al. describe offline licensing, including its claims and threats S-0057 S-0006. Ma et al. describe microarchitectural throttles with a stated adversary S-0036. Petrie describes a license-gated off-switch with a threat model up to states with physical access S-3160. The flexHEG reports describe license-gated operation S-1204.
  • R2 not met. The quantitative evidence is one cycle-accurate simulation of a modelled A100, at the level of kernels rather than end-to-end workloads S-0036. The paper announces no code or hardware release; code is not required, but the results are not end-to-end. Petrie's public proof of concept gates an example Int8 adder in simulation. It leaves out the redundant blocks and constant-time arithmetic that the design's threat model relies on, so it does not test resistance to the attacks those features address S-3161. The existing features that sources mention, thermal and power clock throttling and Intel On Demand feature licensing, were not built or evaluated for this use S-0057 S-0006.

Evidence needed for the next level

  • A public hardware or FPGA prototype, or end-to-end results described in enough detail to repeat (released simulator changes and scripts would also serve).

  • End-to-end evaluation on training and inference workloads on current architectures.

  • A trigger or licensing path that is cryptographically validated and shown to resist interception.

  • Attestation by which a remote verifier can confirm the throttle state.

Limitations, flaws, and blockers

These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.

significant / open / open question

Trigger and licensing path security is unresolved

Ma et al. assume the trigger path stays within the same trust domain as the throttle. They note that an externally sourced trigger from a trusted authority may be intercepted or blocked by the model or by adversarial humans. They state that the trigger must be cryptographically validated and point to prior kill-switch work, but they treat detection and triggering as separate from the throttles they evaluate S-0036. O'Gara et al. list the scalable verification of license authenticity, and how licenses should be issued, as open questions S-0006.

S-0036S-0006

significant / open / theoretical argument

Physical and firmware attacks on the enforcing hardware

RAND's threat analysis includes invasive and semi-invasive physical attacks, fault injection, and firmware and supply-chain attacks. It judges that anti-tamper measures would not be insurmountable for a determined and well-resourced adversary S-0057. Ma et al. assume that neither the model nor human attackers can alter on-chip hardware S-0036. Petrie's design aims to raise the cost of invasive editing by spreading thousands of security blocks through the chip, so that an attack would need thousands of precise edits per chip S-3160.

S-0057S-0036S-3160

minor / open / open question

Sensitivity is architecture-specific and some knobs behave non-monotonically

The authors caution that exact sensitivity curves and knob rankings "may shift across configurations". L2 associativity throttling showed non-monotonic performance because of address-mapping effects S-0036.

S-0036

minor / open / theoretical argument

Workloads can adapt to a throttled resource

A throttled AI could switch to a simpler model, which the authors call "precisely the intended effect". They argue that highly optimized kernels leave little headroom for further adaptation S-0036.

S-0036

What still blocks use or stronger assurance

  1. The throttles need new microarchitecture in future chips, and chipmakers would have to adopt it.

    S-0036S-0056S-3162
  2. Secure licensing and trigger infrastructure is missing, such as a guarantee processor that issues or checks licenses.

    Dependency: Hardware-enabled guarantees (flexHEG) and guarantee processors

    S-1204S-0036
  3. Licenses denominated in work need secure meters for the licensed quantities.

    Dependency: On-chip telemetry from timing, memory and performance counters

    S-0006
  4. No throttle has been evaluated on real hardware or against red-team attempts at bypass.

    S-0036

Connections in the research map

Depends on

Complementary techniques

Concepts used

Organizations and developers

The Consortium’s case files

Related editorial reviews use the Consortium’s own descriptive scores and review dates. Their scores are separate from the atlas readiness rubric.

OS-00 / ControlTen thousand off-switchesRead case file ↗OL-02 / ControlRent-to-run siliconRead case file ↗

Sources and provenance

  1. S-0036 / Tier B

    Hardware Mechanisms to Dynamically Throttle AI Performance ↗

    H. Ma, J. Forzani, L. Malek, D. Wentzlaff · 2026 · arXiv

    Supports: microarchitectural throttling knobs, threat model, simulation setup and results, limitations

    Locator: Abstract; §3-§6

    Version and catalogue details
  2. S-0057 / Tier B

    Hardware-Enabled Governance Mechanisms: Developing Technical Solutions to Exempt Items Otherwise Classified Under Export Control Classification Numbers 3A090 and 4A090 ↗

    G. Kulp, D. Gonzales, E. Smith, L. Heim, P. Puri, M. J. D. Vermeer, Z. Winkelman · 2024 · RAND Corporation

    Supports: offline licensing design and example, existing clock throttling, threat actors, anti-tamper limits

    Locator: pp. viii-ix, 19-27

    Version and catalogue details
  3. S-0006 / Tier B

    Hardware-Enabled Mechanisms for Verifying Responsible AI Development ↗

    A. O'Gara, G. Kulp, W. Hodgkins, J. Petrie, V. Immler, A. Aysu, K. Basu, S. Bhasin, S. Picek, A. Srivastava · 2025 · arXiv

    Supports: licenses as keys for a set amount of work, throttling actions, meters, Intel On Demand, open questions

    Locator: §2.5, §2.5.2, §2.5.4

    Version and catalogue details
  4. S-0056 / Tier B

    Secure, Governable Chips: Using On-Chip Mechanisms to Manage National Security Risks from AI & Advanced Computing ↗

    O. Aarne, T. Fist, C. Withers · 2024 · Center for a New American Security

    Supports: operating licenses in a hardened security module; effort estimate for adequate hardware security

    Locator: Key findings

    Version and catalogue details
  5. S-0035 / Tier B

    Flexible Hardware-Enabled Guarantees for AI Compute ↗

    J. Petrie, O. Aarne, N. Ammann, D. Dalrymple · 2025 · arXiv

    Supports: guarantee processor blocking operations above a threshold

    Locator: Executive Summary

    Version and catalogue details
  6. S-1204 / Tier B

    Technical Options for Flexible Hardware-Enabled Guarantees ↗

    J. Petrie, O. Aarne · 2025 · arXiv

    Supports: periodic licenses with minimum version; multi-party update approval

    Locator: section on updates

    Version and catalogue details
  7. S-3160 / Tier B

    Embedded Off-Switches for AI Compute ↗

    J. Petrie · 2025 · arXiv

    Supports: embedded off-switch design, threat actors, per-block transistor and area estimates, redundancy against invasive edits; no built hardware

    Locator: Abstract; threat model; overhead and attack sections

    Version and catalogue details
  8. S-3161 / Tier B

    JamesPetrie/off-switch (GitHub repository) ↗

    J. Petrie · 2025 · GitHub

    Supports: SystemVerilog proof of concept gating an example Int8 adder in simulation; signature-check cycle counts and omitted production features (README)

    Locator: README

    Version and catalogue details
  9. S-3162 / Tier B

    No Backdoors. No Kill Switches. No Spyware. ↗

    D. Reber Jr. · 2025 · NVIDIA Blog

    Supports: NVIDIA's stated position that its GPUs do not and should not have kill switches, and its distinction for optional user-controlled features

    Locator: blog post

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review