- How many meter-defined events the covered accelerator observed.
- Whether that local counter crossed a configured numerical threshold.
- Which attested device produced a reported count, assuming the key and meter remain intact.
Mechanism under review
FLOPs with a notary stamp
On-chip metering would turn accelerator operations into authenticated counts that can support reporting, licences or automatic caps. The source compares distributed security blocks, an enclosed guarantee processor and adapted performance counters, and rates all governance-grade variants as R&D. Even a perfect counter measures events under a chosen accounting rule; capability, risk, run boundaries and compute performed elsewhere remain policy and system questions.
Read primary source ↗Explicitly tiered
The source distinguishes three attacker tiers and directly addresses physical attacks, structuring, distributed training and algorithmic substitution.Cross-layer
Measurement semantics, silicon integrity, authentication, aggregation and threshold policy must all compose before the count supports compliance.Treaty-scale
The source positions metering as a foundation for domestic thresholds, licences and multilateral treaty verification.Produce trustworthy per-chip compute measurements for threshold-based reporting, verification or enforcement.
- 01
Observe selected arithmetic instructions or data transfers using embedded blocks, an auxiliary guarantee processor or adapted performance counters.
- 02
Maintain a tamper-resistant running total or decrementing usage allowance across resets and power loss.
- 03
Bind the measurement to a device identity and authenticate it for a verifier.
- 04
Aggregate covered measurements into a training-run record, or trigger licensing and halt logic at a configured threshold.
An authenticated count, allowance state or threshold event for operations visible to the covered meter.
What the primitive says—and what it does not.
- Whether the events belong to one legally defined training run across chips, providers and time.
- Whether counted FLOPs correspond to model capability, risk or regulatory intent.
- Whether distillation, fine-tuning, inference-time scaling or algorithmic efficiency moves the activity outside the chosen boundary.
- Whether undeclared or unequipped compute completed part of the workload.
The event tap, counter state, device key, aggregation rule and any licence or halt path driven by the count.
Commercial, well-resourced non-state and nation-state operators, including actors with physical possession and incentives to structure work around thresholds.
The mechanism observes only events exposed to meters on equipped hardware; run aggregation and the legal meaning of the count are outside the counter.
Capabilities considered
- Modify firmware, reset exposed counters or exploit licence logic.
- Use side channels, voltage or laser glitching, probing and focused-ion-beam editing.
- Split training across providers, chips or declared runs and disguise communication patterns.
- Gain algorithmic efficiency, distil a larger model or move consequential work to fine-tuning and inference.
Limits and exclusions
- The taxonomy reports public designs rather than a functional prototype of its own.
- The performance cost of mediating all I/O through the enclosed guarantee processor is unresolved.
- The source treats capability- or output-based evaluation as a needed supplement, not something the meter supplies.
The assurance dependency chain.
The selected hardware events must implement a complete, non-duplicative FLOP-accounting standard.
The meter is precise but counts the wrong policy object.
centralCounter state and event taps must resist reset, bypass, replay and physical manipulation.
Authentic reports contain attacker-controlled totals.
centralAttestation keys and the reporting path must bind readings to the intended accelerator.
A verifier receives substituted or forged measurements.
discussedCounts across chips, providers and time must be assigned to the correct regulated activity.
Structuring leaves every local report compliant while aggregate compute exceeds the threshold.
discussedThe numerical threshold must continue to track the capability or risk that policy intends to govern.
Efficiency and lifecycle changes make the threshold obsolete or overinclusive.
centralLimitations the source already recognizes.
- On-chip metering is rated 'requires R&D'; no functional prototype exists and deployment would lag development by years (§3.1, pp. 5–6).
- Distributed blocks expose many physical targets, the enclosed processor creates a single point of failure, and adapted counters are weakest against adversaries (§3.1, pp. 5–6).
- Algorithmic efficiency, distillation and ambiguous model boundaries can make static compute thresholds obsolete (§4.1, pp. 10–12).
- Training can be structured across providers or decentralised over lower-bandwidth links to evade aggregation and reporting (§4.2, p. 12).
- The proposed security standard is cost and tamper evidence rather than absolute prevention against nation-state custody (§4.3, pp. 12–13).
A protected counter can strengthen a factual claim about covered operations. The overreach begins when that count is treated as a self-interpreting measure of a model, a capability, a training run or compliance across an ecosystem.
Load-bearing sequence
- The event definition matches the accounting standard.
- Every relevant operation reaches an intact meter exactly as intended.
- Authenticated local counts are completely aggregated into the regulated run.
- The policy threshold remains meaningful despite technical change.
Institutional translationThe arithmetic has acquired a notary; the definition of the transaction remains with counsel.
A finding should be falsifiable.
Compare meter totals against independent traces across mixed precision, sparsity, recomputation and fused kernels.
Attempt reset, rollback, replay, event-tap bypass and fault injection under each proposed architecture.
Split a training objective across chips, providers, runs, fine-tuning and distillation and test whether aggregation reconstructs it.
Measure guarantee-processor throughput and latency while mediating worst-case accelerator I/O.
Backtest candidate thresholds against models with similar capabilities but substantially different training compute.
Evidence register (4)
Three metering architectures, feasibility, overhead, counter adaptation and architecture-specific weaknesses.
Threshold obsolescence, accounting ambiguity, distillation and decentralised training.
Adversary tiers, physical attack capabilities and tamper-evident security standard.
Use of on-chip metering in domestic, bilateral and treaty scenarios.