K-0020

Sampling and assurance

Checking a random sample of accelerators, workload segments or outputs rather than all of them, so that violations are caught with a calculable probability.

Source reviewed 2026-09-25

01 / The mechanism and its boundary

What is being described

Sampling-based assurance checks a random sample of accelerators, workload segments or outputs instead of all of them, and chooses the sample size so that a violation is caught with a desired probability S-0029.

Shavit gives a formula for how many accelerators a verifier must inspect in each monitoring period to find at least one accelerator used in a rule-violating training run with a chosen probability S-0029. The required number falls as the run occupies a larger share of the prover's accelerators, so larger runs need fewer inspections S-0029. Sampling works only if the prover cannot predict what will be checked S-0029 or change its records once it knows. In one scheme, the prover commits a hash of sampled weights at each training step before it learns whether that step will be audited S-0017. The same logic applies to recomputation of random training segments in proof-of-learning S-0029 and of random workload samples in reproducible computation packets S-0067. Against a covert adversary, sampling works through deterrence: one system overview notes that such an adversary is caught if it fails to stay hidden even once, and expects physical security and randomly sampled inspections to be the primary defences S-0018.

Connections in the research map

Related research

Sources and provenance

  1. S-0029 / Tier B

    What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring ↗

    Y. Shavit · 2023 · arXiv

    Supports: number of chips to sample per monitoring period to catch at least one chip from a violating run with probability p; fewer samples for larger runs; the Prover cannot predict which chips are inspected; sampled segment recomputation

    Locator: §3.1–3.2, Equation 1, Table 1; §5.1

    Version and catalogue details
  2. S-0017 / Tier C

    Example Schemes for Verifying High-Stakes AI Agreements ↗

    Amodo Design · 2026 · Amodo Design

    Supports: hash commitment to sampled weights before the prover learns whether a step will be audited

    Locator: pre-training scheme

    Version and catalogue details
  3. S-0067 / Tier C

    Verification Plan ↗

    R. Dean · 2026 · AI 2040

    Supports: recomputation server re-runs random samples of workload packets

    Locator: Concrete inference-only retrofitting proposal

    Version and catalogue details
  4. S-0018 / Tier B

    A System Overview for Near-Term, Low-Trust AI Compute Verification ↗

    N. Cankaya · 2026 · Machine Intelligence Research Institute

    Supports: a covert adversary is caught if it fails to stay hidden even once; random sampling needs to catch only a single instance of cheating; physical security and randomly sampled inspections as the likely primary defences

    Locator: threat model; defence layers

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review