M-0002 / Cryptography & computation

Deterministic and bit-exact inference

Making model inference reproducible bit for bit, so that a verifier's re-run must match the provider's output exactly rather than approximately.

R3 In productionSource reviewed 2026-09-25Provider-reported evidence

01 / The mechanism and its boundary

What the technique establishes

Deterministic inference makes an AI model's outputs reproducible bit for bit, so a verifier's re-run must match exactly. Otherwise, re-runs on the same input often differ slightly, because floating-point results depend on the order of operations, which shifts with batch size, hardware and software. This noise forces recomputation checks to accept approximate matches, which a cheating provider could exploit. Three methods remove it: kernels whose results do not depend on batch size, recording the hardware and software needed to reproduce rounding errors, or reproducible operators that fix rounding across hardware. vLLM and SGLang offer deterministic modes, and Gensyn and Eigen Labs report verification services built on exact replay. A published emulator reproduces dense transformer blocks bit for bit on four NVIDIA GPU models. The obstacles are the throughput cost of batch-invariant kernels, gaps in the emulator's coverage, no independent security evaluation, and the provider's need to disclose its full configuration.

Threat model
Adversarial prover
Adversarial evaluation
Published analysis
Hardware needed
None
Prover cooperation
Required
Confidentiality
Partial
Category
Cryptography & computation
Technical detail and cited results

Three routes lead to exact results.

  • Invariance. Kernels fix the reduction order for each output element regardless of batch size. Thinking Machines made RMSNorm, matrix multiplication and attention batch-invariant, the last with a fixed split size for the key-value dimension rather than a fixed number of splits S-1009. vLLM exposes this behind VLLM_BATCH_INVARIANT=1 on NVIDIA GPUs of compute capability 8.0 or higher and on Intel XPUs, in beta S-1013. SGLang integrated batch-invariant attention for its FlashInfer, FlashAttention 3 and Triton backends S-1012. LLM-42 instead decodes on a non-deterministic fast path and replays candidate tokens under a fixed-shape reduction schedule, rolling back any that are inconsistent S-1014.

  • Record and replay. Stock engines are deterministic but not invariant. Outputs are bitwise reproducible if the verifier knows the hardware model, the exact deployed weights, the parallelism topology (separately for prefill and decode), software versions including custom kernels, and the batch size of each forward pass S-0018 S-0020. Of these, only batch size changes during serving, and it costs one extra integer per forward pass to record S-0020. A software emulator reproduces the rounding of other GPU models by modelling tensor-core accumulation and kernel-specific reduction trees S-0020. Hawkeye reproduces tensor-core matrix multiplication exactly on a CPU for Ampere, Hopper and Ada Lovelace GPUs in FP16, BF16 and FP8 S-1010.

  • Reproducible operators. Gensyn's RepOps fixes the order of floating-point operations across hardware, so providers and a referee can reproduce the same result S-1809.

With exact replay, verification is pass/fail, and the chance of catching at least one false output in k samples is 1 − (1 − p)^k for a false-output rate p S-0020.

Claims and scope

A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.

Readiness for a stated use

R3 In production

Assessed use: reproducing open-model inference from receipts in Gensyn's information-market service

low confidence · current · assessed 2026-10-08 · rubric 1.1

This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.

Gensyn reports using exact replay in production so that anyone can check markets it settles with REE, but the production evidence comes from Gensyn, and no independent security evaluation of bit-exact verification has been published.

  • R1 met: the verification claim, a covert-adversary threat model and the information a verifier needs are published S-0020.
  • R2 met: public code predicts dense LLM blocks bit for bit on A100, L40, L40S and H100 GPUs running unmodified vLLM and Hugging Face engines, against a stated adversary S-0020. Batch-invariant modes are public in vLLM (beta) and SGLang S-1013 S-1012, and batch invariance was shown on a 235-billion-parameter model S-1009. Exact CPU reproduction of GPU matrix multiplication is peer-reviewed S-1010.
  • R3 met through Gensyn's REE, on its developer's account. Gensyn reports that Delphi, its information-market app, is live on its mainnet, and that markets settled by open models inside REE produce receipts that anyone can re-run to verify the answer S-3022. Its service documentation also describes the REE judge receipts and mainnet deployment S-0075. REE is publicly available, and its README lists v0.8.0, released on 5 October 2026 S-1812. Gensyn does not label its releases alpha or beta S-3023. Eigen Labs' service does not count towards R3, because Eigen Labs launched it as a mainnet alpha S-3021. No party other than a developer is documented relying on exact replay for a verification decision. The evidence for replay of stock serving engines remains at R2 and covers the emulator's tested dense model blocks S-0020.
  • R4 not met: as of September 2026 no independent audit, red-team or peer-reviewed security analysis of bit-exact verification has been published.

Confidence is low: R3 rests on Gensyn's account of its own service.

Evidence needed for the next level

  • An independent public evaluation (audit, red-team or peer-reviewed security analysis) of bit-exact verification that leaves no critical flaw open.

Limitations, flaws, and blockers

These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.

minor / open / open question

Some kernels remain genuinely nondeterministic

The bit-exact work separates kernels that are deterministic but not batch-invariant from truly nondeterministic ones that use atomic functions. Some integer de-quantization kernels use atomic additions and remain nondeterministic, so exact replay needs backends that avoid them S-0020.

S-0020

significant / open / open question

Cross-hardware replay relies on reverse-engineered, closed behaviour

Emulating one GPU's rounding on another requires reverse-engineering tensor-core arithmetic and modelling proprietary kernel choices. Hawkeye covers a subset of NVIDIA architectures and states that attention and other higher-level operations need further reverse engineering S-1010. For the bit-exact emulator, a proprietary Hopper kernel family is an open edge case S-0020.

S-1010S-0020

What still blocks use or stronger assurance

  1. Batch-invariant kernels cost throughput: in Thinking Machines' Qwen3-8B test, an improved deterministic build took 42 s against 26 s for vLLM's default, and SGLang reports an average 34.35% slowdown on its FlashInfer and FlashAttention 3 backends.

    S-1009S-1012
  2. Coverage is incomplete: the bit-exact emulator targets dense blocks on NVIDIA GPUs and excludes mixture-of-experts inference and training, and vLLM's batch-invariant mode is in beta, with open work on AMD hardware and speculative decoding.

    S-0020S-1013S-1814
  3. Amodo's status page for the AI 2040 verification plan rates a reproducible inference stack for that plan as 'not started'.

    S-1008
  4. Exact replay requires the prover to disclose weights, software versions, parallelism and batch sizes to whoever recomputes.

    S-0020S-0018

Connections in the research map

Complementary techniques

Concepts used

Implementations

Sources and provenance

  1. S-0020 / Tier B

    Bit-Exact AI Inference Verification Without Performance Tradeoffs ↗

    N. Cankaya · 2026 · ICML 2026 Workshop on Technical AI Governance Research

    Supports: covert-adversary threat model; deterministic but non-invariant engines; required metadata; software emulator and its results; limitations; comparison with statistical schemes; best-paper award and public code (arXiv comments)

    Locator: abstract; §1; results; limitations section; arXiv comments

    Version and catalogue details
  2. S-0015 / Tier B

    Verifying LLM Inference to Detect Model Weight Exfiltration ↗

    R. Rinberg, A. Karvonen, A. Hoover, D. Reuter, K. Warr · 2025 · arXiv

    Supports: logging of inferences and random sampling for verification as components separate from recomputation

    Locator: §4.2, assumptions 2 and 3 (v3)

    Version and catalogue details
  3. S-0018 / Tier B

    A System Overview for Near-Term, Low-Trust AI Compute Verification ↗

    N. Cankaya · 2026 · Machine Intelligence Research Institute

    Supports: replay metadata in a low-trust verification system

    Locator: §5.2.2

    Version and catalogue details
  4. S-0016 / Tier B

    DiFR: Inference Verification Despite Nondeterminism ↗

    A. Karvonen, D. Reuter, R. Rinberg, L. Marks, A. Garriga-Alonso, K. Warr · 2025 · ICML 2026 Workshop on Technical AI Governance Research

    Supports: sources of benign nondeterminism; fixed-hardware determinism vs heterogeneous deployments; >98% token agreement

    Locator: §2; §3; §7

    Version and catalogue details
  5. S-0017 / Tier C

    Example Schemes for Verifying High-Stakes AI Agreements ↗

    Amodo Design · 2026 · Amodo Design

    Supports: floating-point non-associativity; fuzzy comparison in recomputation schemes

    Locator: determinism discussion

    Version and catalogue details
  6. S-0067 / Tier C

    Verification Plan ↗

    R. Dean · 2026 · AI 2040

    Supports: reproducibility required for packet correctness checks

    Locator: Concrete inference-only retrofitting proposal

    Version and catalogue details
  7. S-1008 / Tier C

    AI 2040 Plan A — Verification SITREP ↗

    Amodo Design · 2026 · Amodo Design

    Supports: status of reproducible inference stack

    Locator: status items

    Version and catalogue details
  8. S-1009 / Tier C

    Defeating Nondeterminism in LLM Inference ↗

    H. He, Thinking Machines Lab · 2025 · Thinking Machines Lab: Connectionism

    Supports: batch invariance as main cause; kernels made invariant; Qwen3-235B experiment; timings

    Locator: whole post

    Version and catalogue details
  9. S-1010 / Tier A

    Hawkeye: Reproducing GPU-Level Non-Determinism ↗

    E. Badash, D. Boneh, I. Komargodski, M. Srivastava · 2026 · Proceedings of Machine Learning and Systems 8 (MLSys 2026)

    Supports: exact CPU reproduction of tensor-core matrix multiplication; scope limits

    Locator: abstract; §8; §9

    Version and catalogue details
  10. S-1011 / Tier B

    DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence ↗

    DeepSeek-AI · 2026 · arXiv

    Supports: provider-reported end-to-end batch-invariant and deterministic kernels

    Locator: §3.3

    Version and catalogue details
  11. S-1012 / Tier C

    Towards Deterministic Inference in SGLang and Reproducible RL Training ↗

    The SGLang Team · 2025 · LMSYS Org blog

    Supports: SGLang deterministic mode and its overhead

    Locator: whole post

    Version and catalogue details
  12. S-1013 / Tier B

    Batch Invariance (vLLM documentation) ↗

    vLLM project · 2026 · vLLM documentation (GitHub, docs/features/batch_invariance.md)

    Supports: vLLM batch-invariance flag, supported hardware (NVIDIA compute capability 8.0+, Intel XPUs), beta status

    Locator: whole page

    Version and catalogue details
  13. S-1014 / Tier B

    LLM-42: Enabling Determinism in LLM Inference with Verified Speculation ↗

    R. Gond, A. K. Kamath, R. Ramjee, A. Panwar · 2026 · arXiv

    Supports: scheduling-based determinism alternative that mostly reuses existing kernels

    Locator: abstract

    Version and catalogue details
  14. S-1813 / Tier B

    thinking-machines-lab/batch_invariant_ops (GitHub repository) ↗

    Thinking Machines Lab · 2025 · GitHub

    Supports: Thinking Machines' batch-invariant kernel library (MIT)

    Locator: README

    Version and catalogue details
  15. S-1814 / Tier C

    [Feature]: Batch Invariant Feature and Performance Optimization (vLLM issue #27433) ↗

    vLLM project contributors · 2025 · GitHub (vllm-project/vllm issues)

    Supports: vLLM developers' statement that batch-invariance support is based on the Thinking Machines post; open work on AMD hardware, NVFP4 and speculative decoding

    Version and catalogue details
  16. S-1809 / Tier B

    Verde: Verification via Refereed Delegation for Machine Learning Programs ↗

    A. Arun, A. St. Arnaud, A. Titov, B. Wilcox, V. Kolobaric, M. Brinkmann, O. Ersoy, B. Fielding, J. Bonneau · 2025 · arXiv

    Supports: Verde refereed delegation and RepOps reproducible operators

    Locator: abstract; §3.2

    Version and catalogue details
  17. S-1810 / Tier C

    Verde Verification System In Production ↗

    O. Ersoy · 2025 · Gensyn research blog

    Supports: Gensyn's report of Verde and RepOps in production (provider-reported)

    Version and catalogue details
  18. S-1812 / Tier B

    gensyn-ai/ree: Gensyn Reproducible Execution Environment (GitHub repository) ↗

    Gensyn · 2026 · GitHub

    Supports: REE public as binaries with an MIT-licensed SDK; v0.8.0 released on 5 October 2026 (provider-reported)

    Locator: README; patch notes

    Version and catalogue details
  19. S-3020 / Tier B

    EigenAI: Deterministic Inference, Verifiable Results ↗

    D. Ribeiro Alves, V. Patankar, M. Pereira, J. Stephens, N. Vaziri, S. Kannan · 2026 · arXiv

    Supports: EigenAI deterministic engine and optimistic re-execution protocol; same-SKU and A100-vs-H100 determinism results (provider-reported)

    Locator: abstract; Table 5

    Version and catalogue details
  20. S-3021 / Tier C

    EigenCloud Brings Verifiable AI to Mass Market with EigenAI and EigenCompute Launches ↗

    EigenCloud · 2025 · Eigen Labs blog

    Supports: EigenAI mainnet alpha launch; stake not yet exposed to slashing (provider-reported)

    Version and catalogue details
  21. S-3022 / Tier C

    Building Delphi: Pricing, Settlement, and Agentic Trading ↗

    D. Jedamski · 2026 · Gensyn blog

    Supports: Delphi live on Gensyn's mainnet; REE settlement receipts that anyone can re-run (provider-reported)

    Locator: settlement section

    Version and catalogue details
  22. S-3023 / Tier B

    Reproducible Execution Environment (REE) (Gensyn documentation) ↗

    Gensyn · 2026 · Gensyn documentation

    Supports: REE releases carry no alpha or beta label (provider-reported)

    Locator: whole page

    Version and catalogue details
  23. S-0075 / Tier B

    What is Delphi? (Delphi documentation) ↗

    Gensyn · 2026 · Delphi documentation

    Supports: Delphi mainnet service and REE judge receipts that anyone can re-run (provider-reported)

    Locator: Settled by AI, Verifiable by Anyone; Where Delphi Runs

    Version and catalogue details
  24. S-3500 / Tier A

    TAO: Tolerance-Aware Optimistic Verification for Floating-Point Neural Networks ↗

    J. Yao, H. Su, T. Liao, Z. Cheng, H. Zhang, X. Wang, P. Viswanath · 2026 · Proceedings of the 21st European Conference on Computer Systems (EuroSys 2026), pp. 1515-1532

    Supports: TAO: operator-level tolerance bounds instead of bitwise equality on heterogeneous hardware; Merkle-anchored dispute game; Ethereum testnet deployment

    Locator: abstract; §1; evaluation

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review