01 / The mechanism and its boundary
What the technique establishes
Throttling mechanisms would let the hardware itself enforce limits on AI computation. One proposal is offline licensing: a chip runs at full speed only while it holds a valid license for a set amount of work. Another is a set of microarchitectural throttles that, when triggered, shrink L2 cache capacity, add cache latency, cap cache bandwidth or limit shared-memory access. A 2026 simulation of an NVIDIA A100 found that such throttles could cut the performance of large-language-model kernels by up to 80% at one-eighth of resource availability, at a cost of under about 10,000 flip-flops each. The results come from simulation, not real chips. For verification, throttling matters only if a verifier can confirm the limit is in place and cannot be bypassed. That depends on attestation, tamper resistance and a secure trigger or licensing path. RAND judges that anti-tamper protection would not be insurmountable for a determined, well-resourced adversary.
- Threat model
- Adversarial prover
- Adversarial evaluation
- Published analysis
- Hardware needed
- New chip design
- Prover cooperation
- Partial
- Confidentiality
- Preserving
- Category
- On chip & hardware
Technical detail and cited results
- Knobs evaluated by Ma et al. L2 capacity through way masking, L2 latency through configurable request buffering, L2 bandwidth through credit-based rate limiting, and shared-memory port access through bank arbitration. None is exposed to a software-visible interface S-0036.
- Setup. The simulator was AccelSim, configured as an NVIDIA A100, with utilization metrics validated on a real GPU with Nsight Compute. The workloads were CUTLASS StreamK GEMM kernels, for prefill (M=4096) and decode (M=128), and FlashAttention, with shapes from DeepSeek-V3, Llama-3-70B and Mixtral-8x7B, plus six non-LLM CUDA sample workloads S-0036.
- Results. At one-eighth of resource availability, performance fell by up to 80% (L2 latency for decode; L2 associativity for prefill). After a throttle was applied, performance settled within about 5K cycles for shared-memory ports, 5–7K cycles for L2 response rate and about 80K cycles for L2 associativity. Each mechanism needs fewer than about 10K flip-flops S-0036.
- Selectivity. Memory-side knobs affected LLM kernels selectively. Disabling compute cores degraded nearly every workload together S-0036.
- Offline licensing (RAND). A secure message authenticated on the GPU might authorize, for example, 10^18 arithmetic operations, after which the GPU would fall back to 1% of full performance S-0057.
- Embedded off-switch (Petrie). Each security block uses public-key signatures, with nonces against replay, to check for recent authorization. Petrie estimates roughly 40,000 transistors per block, and about 0.5% of the die for 10,000 blocks S-3160. The proof-of-concept README reports signature checks of about 5 million cycles for ECDSA and about 300,000 cycles for HSS/LMS S-3161.
Claims and scope
A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.
Compute stock is at most a declared amount
A verified performance cap bounds the effective capacity of declared hardware; needs attestation that the cap is active.
A training run stayed within declared limits
Licenses that authorize a fixed amount of work would bound compute per license period (S-0057, S-0006).
Declared hardware is idle or shut down
Unlicensed hardware falls back to reduced capacity or shuts down (S-0006).
Readiness for a stated use
Assessed use: performance limits a verifier can rely on, against an operator trying to bypass them
medium confidence · current · assessed 2026-09-25 · rubric 1.1
This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.
The designs are published, but the only quantitative results are kernel-level simulations and the only public code is a simulated proof of concept.
- R1 met: RAND and O'Gara et al. describe offline licensing, including its claims and threats S-0057 S-0006. Ma et al. describe microarchitectural throttles with a stated adversary S-0036. Petrie describes a license-gated off-switch with a threat model up to states with physical access S-3160. The flexHEG reports describe license-gated operation S-1204.
- R2 not met. The quantitative evidence is one cycle-accurate simulation of a modelled A100, at the level of kernels rather than end-to-end workloads S-0036. The paper announces no code or hardware release; code is not required, but the results are not end-to-end. Petrie's public proof of concept gates an example Int8 adder in simulation. It leaves out the redundant blocks and constant-time arithmetic that the design's threat model relies on, so it does not test resistance to the attacks those features address S-3161. The existing features that sources mention, thermal and power clock throttling and Intel On Demand feature licensing, were not built or evaluated for this use S-0057 S-0006.
Evidence needed for the next level
A public hardware or FPGA prototype, or end-to-end results described in enough detail to repeat (released simulator changes and scripts would also serve).
End-to-end evaluation on training and inference workloads on current architectures.
A trigger or licensing path that is cryptographically validated and shown to resist interception.
Attestation by which a remote verifier can confirm the throttle state.
Limitations, flaws, and blockers
These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.
significant / open / open question
Trigger and licensing path security is unresolved
Ma et al. assume the trigger path stays within the same trust domain as the throttle. They note that an externally sourced trigger from a trusted authority may be intercepted or blocked by the model or by adversarial humans. They state that the trigger must be cryptographically validated and point to prior kill-switch work, but they treat detection and triggering as separate from the throttles they evaluate S-0036. O'Gara et al. list the scalable verification of license authenticity, and how licenses should be issued, as open questions S-0006.
significant / open / theoretical argument
Physical and firmware attacks on the enforcing hardware
RAND's threat analysis includes invasive and semi-invasive physical attacks, fault injection, and firmware and supply-chain attacks. It judges that anti-tamper measures would not be insurmountable for a determined and well-resourced adversary S-0057. Ma et al. assume that neither the model nor human attackers can alter on-chip hardware S-0036. Petrie's design aims to raise the cost of invasive editing by spreading thousands of security blocks through the chip, so that an attack would need thousands of precise edits per chip S-3160.
minor / open / open question
Sensitivity is architecture-specific and some knobs behave non-monotonically
The authors caution that exact sensitivity curves and knob rankings "may shift across configurations". L2 associativity throttling showed non-monotonic performance because of address-mapping effects S-0036.
minor / open / theoretical argument
Workloads can adapt to a throttled resource
A throttled AI could switch to a simpler model, which the authors call "precisely the intended effect". They argue that highly optimized kernels leave little headroom for further adaptation S-0036.
What still blocks use or stronger assurance
- S-0036S-0056S-3162
The throttles need new microarchitecture in future chips, and chipmakers would have to adopt it.
Secure licensing and trigger infrastructure is missing, such as a guarantee processor that issues or checks licenses.
Dependency: Hardware-enabled guarantees (flexHEG) and guarantee processors
S-1204S-0036Licenses denominated in work need secure meters for the licensed quantities.
Dependency: On-chip telemetry from timing, memory and performance counters
S-0006- S-0036
No throttle has been evaluated on real hardware or against red-team attempts at bypass.
Connections in the research map
Depends on
- On-chip telemetry from timing, memory and performance counters
Licenses denominated in work need secure meters for the licensed quantities (S-0006).
- TEE remote attestation for AI workloads
A verifier needs attested configuration to know a cap is in force; licensing relies on secure boot and on-chip authentication (S-0057).
Complementary techniques
Organizations and developers
The Consortium’s case files
Related editorial reviews use the Consortium’s own descriptive scores and review dates. Their scores are separate from the atlas readiness rubric.
OS-00 / ControlTen thousand off-switchesRead case file ↗OL-02 / ControlRent-to-run siliconRead case file ↗Sources and provenance
- S-0036 / Tier B
Hardware Mechanisms to Dynamically Throttle AI Performance ↗
H. Ma, J. Forzani, L. Malek, D. Wentzlaff · 2026 · arXiv
Supports: microarchitectural throttling knobs, threat model, simulation setup and results, limitations
Locator: Abstract; §3-§6
Version and catalogue details - S-0057 / Tier B
Hardware-Enabled Governance Mechanisms: Developing Technical Solutions to Exempt Items Otherwise Classified Under Export Control Classification Numbers 3A090 and 4A090 ↗
G. Kulp, D. Gonzales, E. Smith, L. Heim, P. Puri, M. J. D. Vermeer, Z. Winkelman · 2024 · RAND Corporation
Supports: offline licensing design and example, existing clock throttling, threat actors, anti-tamper limits
Locator: pp. viii-ix, 19-27
Version and catalogue details - S-0006 / Tier B
Hardware-Enabled Mechanisms for Verifying Responsible AI Development ↗
A. O'Gara, G. Kulp, W. Hodgkins, J. Petrie, V. Immler, A. Aysu, K. Basu, S. Bhasin, S. Picek, A. Srivastava · 2025 · arXiv
Supports: licenses as keys for a set amount of work, throttling actions, meters, Intel On Demand, open questions
Locator: §2.5, §2.5.2, §2.5.4
Version and catalogue details - S-0056 / Tier B
Secure, Governable Chips: Using On-Chip Mechanisms to Manage National Security Risks from AI & Advanced Computing ↗
O. Aarne, T. Fist, C. Withers · 2024 · Center for a New American Security
Supports: operating licenses in a hardened security module; effort estimate for adequate hardware security
Locator: Key findings
Version and catalogue details - S-0035 / Tier B
Flexible Hardware-Enabled Guarantees for AI Compute ↗
J. Petrie, O. Aarne, N. Ammann, D. Dalrymple · 2025 · arXiv
Supports: guarantee processor blocking operations above a threshold
Locator: Executive Summary
Version and catalogue details - S-1204 / Tier B
Technical Options for Flexible Hardware-Enabled Guarantees ↗
J. Petrie, O. Aarne · 2025 · arXiv
Supports: periodic licenses with minimum version; multi-party update approval
Locator: section on updates
Version and catalogue details - S-3160 / Tier B
Embedded Off-Switches for AI Compute ↗
J. Petrie · 2025 · arXiv
Supports: embedded off-switch design, threat actors, per-block transistor and area estimates, redundancy against invasive edits; no built hardware
Locator: Abstract; threat model; overhead and attack sections
Version and catalogue details - S-3161 / Tier B
JamesPetrie/off-switch (GitHub repository) ↗
J. Petrie · 2025 · GitHub
Supports: SystemVerilog proof of concept gating an example Int8 adder in simulation; signature-check cycle counts and omitted production features (README)
Locator: README
Version and catalogue details - S-3162 / Tier B
No Backdoors. No Kill Switches. No Spyware. ↗
D. Reber Jr. · 2025 · NVIDIA Blog
Supports: NVIDIA's stated position that its GPUs do not and should not have kill switches, and its distinction for optional user-controlled features
Locator: blog post
Version and catalogue details
- Source review date
- 2026-09-25
- Drafted by (source map)
- ai
- Review handles (source map)
- codex-review