M-0024 / Isolation & architecture

Bounding unexplained information in outputs

Limits the hidden information a facility's outputs can carry by measuring how much of those outputs the declared computation fails to predict.

R2 DemonstratedSource reviewed 2026-09-25

01 / The mechanism and its boundary

What the technique establishes

Bounding unexplained information measures how much of the information leaving a facility the declared work cannot explain. If a verifier can predict outputs from the declared model and recorded inputs, little room remains to smuggle out model weights or the results of hidden work. One proposed architecture routes all traffic through a verifier-controlled interlock and challenges the operator to show that random outputs follow from compliant computation. As of September 2026 no prototype results have been published. An instance for language-model inference, with public code, cut the information an attacker could hide to under 0.5% on a 30-billion-parameter model, under benign prompts and at a false-positive rate below 0.01%. An independent study showed that an attacker who chooses the prompts roughly doubles the leakage per token. The main obstacles are closing every other channel, including side channels, and tolerating numerical noise without leaving room for a covert channel.

Threat model
Adversarial prover
Adversarial evaluation
Independent red-team
Hardware needed
Retrofit device
Prover cooperation
Required
Confidentiality
Partial
Category
Isolation & architecture
Technical detail and cited results
  • Fixed-seed sampling likelihood. The verifier knows the seed and recomputes the next-token distribution for each logged position, then scores whether each emitted token is plausible under honest sampling. Gumbel-Max and inverse-probability-transform estimators are given, and three theorems bound the rate at which an adversary can embed information while passing the check S-0015.
  • Reported result. On MoE-Qwen-30B under benign prompt traffic, exfiltratable information fell to under 0.5% at a false-positive rate below 0.01%, a slowdown of more than 200 times S-0015.
  • Attack result. Across six instruction-tuned models of 1 to 32 billion parameters and three seeds, the strongest prompt attack raised leakage from 0.067–0.119 to 0.145–0.286 bits per token and cut the slowdown from 146–254 times to 60–118 times, because the verifier's admissible-token set grows with the model's output entropy S-1507.
  • Compression bounds. LLM-based arithmetic coding with domain-adapted LoRA adapters is 2 times better than the base model; succinct rewrites reach compression ratios of about 0.03; and a question-asking protocol reaches 0.0006–0.004 S-1508. An earlier proposal reports compression below 1 bit per token by arithmetic coding against a weaker reference model, in an experiment with fine-tuned Pythia models S-1509.

Claims and scope

A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.

Readiness for a stated use

R2 Demonstrated

Assessed use: bounding how much hidden information can leave in checked inference outputs

low confidence · current · assessed 2026-09-25 · rubric 1.1

This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.

R2 through the inference-output instance, which has public code and an independent attack study; the facility-level architecture is still a proposal.

  • R1 met: Petrie and Mühlhäuser publish an architecture, the claim it verifies and its setting, in which neither party trusts the other's hardware S-0019.
  • R2 met through the instance for inference outputs. Rinberg et al. bound the covert information in LLM responses beyond what honest sampling from the declared model explains. They publish code and report results on a 30-billion-parameter mixture-of-experts model against a stated adversary who controls the inference server S-0015. The facility-level architecture remains at R1: its paper outlines protocol details, attacks and prototyping plans, not results S-0019.
  • R3 not met: no party other than the developers relies on the bound for a verification decision, and no production-grade system is available.
  • R4 not met. Its evaluation criterion holds for the inference instance only: an independent researcher attacked it in practice, and the flaw shown widens the bound rather than defeating it, so the evaluation left no critical flaw open S-1507. The facility-level design has not been independently evaluated.

Confidence is low. The R2 evidence covers token outputs of a single inference service, not the general bound over all facility outputs S-0019, and the demonstrated bound degrades when the attacker controls prompts S-1507.

Evidence needed for the next level

  • A prototype of the facility-level architecture (interlock, commitments, challenge-based prediction) with published results.

  • A bound that holds against an adversary who controls the prompt distribution, for example with entropy-calibrated tolerances, evaluated independently.

  • Reliance by a party other than the developer on an unexplained-information bound for a verification decision, or a production-grade deployment.

Limitations, flaws, and blockers

These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.

significant / open / demonstrated attack

Prompt-controlled entropy inflation widens the covert channel

Gumbel-based inference verification tolerates token choices that honest GPU nondeterminism could produce, and the size of that tolerated set grows with the model's output entropy. Kezins, an independent researcher, showed that an adversary who controls the prompt distribution can raise output entropy and roughly double the bits leaked per token. Across six models of 1 to 32 billion parameters, this cut the slowdown from 146–254 times under benign prompts to 60–118 times. Kezins argues that architectures built on the same unexplained-information bound inherit this attack surface, and recommends calibrating tolerances against local token entropy rather than benign traffic.

S-1507S-0015

significant / open / theoretical argument

Information the declared computation explains is not bounded

The bound limits unexplained bits only. Outputs that the declared computation fully explains can still carry valuable information: a compression study notes that an adversary with inference access can extract more proprietary information per bit than naive transmission allows.

S-1508S-0019

significant / open / theoretical argument

Channels other than checked outputs are outside the bound

The inference-verification scheme treats side channels as out of scope. A low-trust system design argues that suppressing physical covert bandwidth below kilobits per second is much more achievable than aiming for zero, and that a malicious device can leak one bit of information by deliberately outputting a wrong result.

S-0015S-0018

significant / open / open question

The facility-level design is untested

The compute-verification architecture is described with protocol details, potential attacks and prototyping plans, but no prototype results have been published.

S-0019

What still blocks use or stronger assurance

  1. The prover's compute must be isolated so that all traffic passes through the verifier's interlock; any unmonitored path voids the bound.

    Dependency: Bandwidth limits and compartmentalization

    S-0019
  2. Physical side channels need separate suppression, and one design treats a low residual bandwidth, rather than zero, as the realistic target.

    Dependency: Side-channel suppression for isolated facilities

    S-0018S-0015
  3. Tolerance for numerical nondeterminism sets the size of the residual channel; bit-exact replay would remove it but needs full hardware and software metadata.

    Dependency: Deterministic and bit-exact inference

    S-1507S-0018
  4. Recomputation over confidential weights and inputs needs a protected setting: prover recomputation in a verifier-controlled enclosure, verifier recomputation in a prover-controlled enclosure, or zero-knowledge proofs.

    S-0019
  5. No prototype of the facility-level architecture exists to red-team.

    S-0019

Connections in the research map

Depends on

Complementary techniques

Concepts used

The Consortium’s case files

Related editorial reviews use the Consortium’s own descriptive scores and review dates. Their scores are separate from the atlas readiness rubric.

IV-10 / WorkloadThe honest twin serverRead case file ↗

Sources and provenance

  1. S-0019 / Tier B

    Verifying AI Compute by Bounding Unexplained Information Exfiltration ↗

    J. Petrie, Y. Mühlhäuser · 2026 · ICML 2026 Workshop on Technical AI Governance Research

    Supports: architecture: isolation, interlock, commitments, challenge-based prediction; principle; three confidentiality options; stage of work

    Locator: abstract (read via the ICML 2026 virtual poster page; the OpenReview PDF was not reachable)

    Version and catalogue details
  2. S-0015 / Tier B

    Verifying LLM Inference to Detect Model Weight Exfiltration ↗

    R. Rinberg, A. Karvonen, A. Hoover, D. Reuter, K. Warr · 2025 · arXiv

    Supports: fixed-seed sampling likelihood; theorems; threat model and assumptions; results; code release; side channels out of scope

    Locator: abstract; §4; §5 (Theorems 5.1-5.3); §6; Appendix E

    Version and catalogue details
  3. S-1507 / Tier B

    Adversarial Entropy Inflation Against Gumbel-Based Inference Verification ↗

    N. Kezins · 2026 · arXiv

    Supports: entropy-inflation attack, results and recommended mitigation; applicability to unexplained-information architectures

    Locator: abstract; introduction; method; conclusion

    Version and catalogue details
  4. S-1508 / Tier B

    Haiku to Opus in Just 10 bits: LLMs Unlock Large Compression Gains ↗

    R. Rinberg, A. M. Carrell, S. Henniger, N. Carlini, K. Warr · 2026 · arXiv

    Supports: compressibility of LLM text; egress limiting rationale; dual-use note

    Locator: abstract; §5.1; §5.3

    Version and catalogue details
  5. S-1509 / Tier C

    Preventing model exfiltration with upload limits ↗

    R. Greenblatt · 2024 · AI Alignment Forum

    Supports: upload limits with compression against a weaker model; below 1 bit per token; assumptions; hidden-distillation route; author's uncertainty and probability estimate

    Locator: whole post

    Version and catalogue details
  6. S-0018 / Tier B

    A System Overview for Near-Term, Low-Trust AI Compute Verification ↗

    N. Cankaya · 2026 · Machine Intelligence Research Institute

    Supports: sanitized egress and one-bit fault leakage; side-channel suppression target; exact-replay metadata

    Locator: §4.3.3; §5.2.2; §5.3.1

    Version and catalogue details
  7. S-1300 / Tier B

    Fingerprinting All AI Cluster I/O Without Mutually Trusted Processors ↗

    N. Cankaya, J. Kryś, J. Ng, L. Marks, F. Krückel · 2026 · arXiv

    Supports: residual covert egress of about 40 Mbit/s for a 200k-GPU inference cluster at about 0.1 bits per token after replay checks

    Locator: §5.2

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review