01 / The mechanism and its boundary
What is being described
Sampling-based assurance checks a random sample of accelerators, workload segments or outputs instead of all of them, and chooses the sample size so that a violation is caught with a desired probability S-0029.
Shavit gives a formula for how many accelerators a verifier must inspect in each monitoring period to find at least one accelerator used in a rule-violating training run with a chosen probability S-0029. The required number falls as the run occupies a larger share of the prover's accelerators, so larger runs need fewer inspections S-0029. Sampling works only if the prover cannot predict what will be checked S-0029 or change its records once it knows. In one scheme, the prover commits a hash of sampled weights at each training step before it learns whether that step will be audited S-0017. The same logic applies to recomputation of random training segments in proof-of-learning S-0029 and of random workload samples in reproducible computation packets S-0067. Against a covert adversary, sampling works through deterrence: one system overview notes that such an adversary is caught if it fails to stay hidden even once, and expects physical security and randomly sampled inspections to be the primary defences S-0018.
Connections in the research map
Related research
Sources and provenance
- S-0029 / Tier B
What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring ↗
Y. Shavit · 2023 · arXiv
Supports: number of chips to sample per monitoring period to catch at least one chip from a violating run with probability p; fewer samples for larger runs; the Prover cannot predict which chips are inspected; sampled segment recomputation
Locator: §3.1–3.2, Equation 1, Table 1; §5.1
Version and catalogue details - S-0017 / Tier C
Example Schemes for Verifying High-Stakes AI Agreements ↗
Amodo Design · 2026 · Amodo Design
Supports: hash commitment to sampled weights before the prover learns whether a step will be audited
Locator: pre-training scheme
Version and catalogue details - S-0067 / Tier C
Verification Plan ↗
R. Dean · 2026 · AI 2040
Supports: recomputation server re-runs random samples of workload packets
Locator: Concrete inference-only retrofitting proposal
Version and catalogue details - S-0018 / Tier B
A System Overview for Near-Term, Low-Trust AI Compute Verification ↗
N. Cankaya · 2026 · Machine Intelligence Research Institute
Supports: a covert adversary is caught if it fails to stay hidden even once; random sampling needs to catch only a single instance of cheating; physical security and randomly sampled inspections as the likely primary defences
Locator: threat model; defence layers
Version and catalogue details
- Source review date
- 2026-09-25
- Drafted by (source map)
- ai
- Review handles (source map)
- codex-review