01 / The mechanism and its boundary
What is being described
A threat model describes the capabilities an attacker is assumed to be able to use against a system, such as the information, computing power and control of the system available to it S-0072.
Its purpose is to identify the threats a design must withstand and to rule others explicitly out of scope, since nearly every security system is vulnerable to a sufficiently dedicated and resourceful attacker S-0072. NIST treats threat modelling as a form of risk assessment that models both the attack and the defence side of a system S-1600. Threat models used in AI verification differ in how far each party is trusted:
- Covert adversary. Shavit models the prover as willing to break the rules only if it expects not to be detected S-0029. The Oxford Martin report uses the same framing for states that seek to demonstrate compliance to each other while also looking for ways to circumvent verification S-0004.
- Mutual distrust. One low-trust system overview assumes nation-state adversaries on both sides, including a verifier that may try to exfiltrate the prover's secrets, and relies on redundant checks across devices that each party trusts unilaterally, instead of a single chain of trust S-0018.
Physical access is a recurring issue: Shavit notes that a prover with unlimited physical access to a chip could undermine its attestation and signed-firmware protections, and proposes physical inspections after the fact to detect such tampering S-0029.
Connections in the research map
Related research
Sources and provenance
- S-0072 / Tier A
Guidelines for Writing RFC Text on Security Considerations (RFC 3552, BCP 72) ↗
E. Rescorla, B. Korver, Internet Architecture Board · 2003 · Internet Engineering Task Force
Supports: threat model describes the capabilities an attacker is assumed to deploy, including information, computing capability and control of the system; purpose is to identify threats of concern and rule others out of scope; nearly every system is vulnerable to a sufficiently dedicated and resourceful attacker
Locator: §3
Version and catalogue details - S-1600 / Tier A
NIST Computer Security Resource Center (CSRC) Glossary ↗
National Institute of Standards and Technology · 2026 · NIST Computer Security Resource Center
Supports: NIST definition of threat modeling as a form of risk assessment modelling attack and defence sides
Locator: term: threat_modeling (NIST SP 800-53 Rev. 5)
Version and catalogue details - S-0029 / Tier B
What does it take to catch a Chinchilla? Verifying Rules on Large-Scale Neural Network Training via Compute Monitoring ↗
Y. Shavit · 2023 · arXiv
Supports: Prover as covert adversary; unlimited physical access undermines attestation and signed firmware; physical inspections after the fact may detect such tampering
Locator: §2; §3.1
Version and catalogue details - S-0004 / Tier B
Verification for International AI Governance ↗
B. Harack, R. F. Trager, A. Reuel, D. Manheim, M. Brundage, O. Aarne, A. Scher, Y. Pan, J. Xiao, K. Loke, S. N. Adan, G. Bas, N. A. Caputo, J. C. Morse, J. Ahuja, I. Duan, J. Egan, B. Bucknall, B. Rosen, R. Araujo, V. Boulanin, R. Lall, F. Barez, S. Alvira, C. Katzke, A. Atamli, A. Awad · 2025 · Oxford Martin AI Governance Initiative
Supports: states seeking to demonstrate compliance while seeking ways to circumvent verification, a framing the report says is sometimes termed the covert adversary
Locator: §1.1.2, fn. 40
Version and catalogue details - S-0018 / Tier B
A System Overview for Near-Term, Low-Trust AI Compute Verification ↗
N. Cankaya · 2026 · Machine Intelligence Research Institute
Supports: nation-state adversaries; no single chain of trust; redundancy across devices trusted unilaterally by each party; malicious verifier
Locator: threat model section
Version and catalogue details
- Source review date
- 2026-09-25
- Drafted by (source map)
- ai
- Review handles (source map)
- codex-review