K-0009

Recomputation

Checking a claimed computation by re-running all of it, or a random sample, on hardware the verifier trusts and comparing the results.

Source reviewed 2026-09-25

01 / The mechanism and its boundary

What is being described

Recomputation checks a claimed computation by re-running it, or a random sample of it, from the same inputs on hardware the verifier trusts, and comparing the results with those reported S-0029 S-0016.

Re-running a large training run in full would be very costly, so in Shavit's framework the verifier re-runs selected segments, starting from a reported checkpoint and applying the reported data batches, and accepts if the result is close to the next reported checkpoint; this is the basis of proof-of-learning and training-transcript verification S-0029. For inference, a trusted reference implementation recomputes what the model should have predicted at each generated token S-0016, as in sampled inference recomputation; the same check can limit how much of a model's weights can be hidden in its responses (Bounding unexplained information in outputs) S-0015. Amodo Design distinguishes correctness, meaning that the workloads run match those declared, from completeness, meaning that every workload is reported, and notes that its recomputation schemes address only correctness S-0017. Recomputation also needs:

  • Access to inputs and weights. This raises confidentiality problems, which Shavit addresses with a jointly trusted, air-gapped cluster S-0029; one low-trust design replays the prover's data in air-gapped auditing environments and checks it against hashes committed earlier S-0018.
  • Commitment before sampling. The prover must fix its records, for example by committing to a hash of sampled weights, before it learns which step will be audited S-0017.
  • A way to handle numerical nondeterminism. Benign noise otherwise makes legitimate variation hard to tell from real problems S-0016.

Connections in the research map

Related research

Sources and provenance

  1. S-0029 / Tier B

    What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring ↗

    Y. Shavit · 2023 · arXiv

    Supports: full re-running of large training is infeasible; verifier re-runs segments from a reported checkpoint with the reported data batches and accepts if close to the next checkpoint; jointly trusted air-gapped cluster for confidentiality

    Locator: §5.1; §5.2

    Version and catalogue details
  2. S-0016 / Tier B

    DiFR: Inference Verification Despite Nondeterminism ↗

    A. Karvonen, D. Reuter, R. Rinberg, L. Marks, A. Garriga-Alonso, K. Warr · 2025 · ICML 2026 Workshop on Technical AI Governance Research

    Supports: trusted reference implementation recomputes predictions for generated tokens; nondeterminism makes legitimate variation hard to tell from real problems

    Locator: abstract

    Version and catalogue details
  3. S-0015 / Tier B

    Verifying LLM Inference to Detect Model Weight Exfiltration ↗

    R. Rinberg, A. Karvonen, A. Hoover, D. Reuter, K. Warr · 2025 · arXiv

    Supports: inference verification against a reference to limit steganographic weight exfiltration

    Locator: abstract

    Version and catalogue details
  4. S-0017 / Tier C

    Example Schemes for Verifying High-Stakes AI Agreements ↗

    Amodo Design · 2026 · Amodo Design

    Supports: prover commits to sampled weights before learning whether a step will be audited; correctness vs completeness; schemes address correctness only

    Locator: pre-training scheme; introduction

    Version and catalogue details
  5. S-0018 / Tier B

    A System Overview for Near-Term, Low-Trust AI Compute Verification ↗

    N. Cankaya · 2026 · Machine Intelligence Research Institute

    Supports: replay of the prover's data in air-gapped auditing environments, checked against committed hashes

    Locator: §1; §5.1; §5.2.1

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review