01 / The mechanism and its boundary
What the technique establishes
Deterministic inference makes an AI model's outputs reproducible bit for bit, so a verifier's re-run must match exactly. Otherwise, re-runs on the same input often differ slightly, because floating-point results depend on the order of operations, which shifts with batch size, hardware and software. This noise forces recomputation checks to accept approximate matches, which a cheating provider could exploit. Three methods remove it: kernels whose results do not depend on batch size, recording the hardware and software needed to reproduce rounding errors, or reproducible operators that fix rounding across hardware. vLLM and SGLang offer deterministic modes, and Gensyn and Eigen Labs report verification services built on exact replay. A published emulator reproduces dense transformer blocks bit for bit on four NVIDIA GPU models. The obstacles are the throughput cost of batch-invariant kernels, gaps in the emulator's coverage, no independent security evaluation, and the provider's need to disclose its full configuration.
- Threat model
- Adversarial prover
- Adversarial evaluation
- Published analysis
- Hardware needed
- None
- Prover cooperation
- Required
- Confidentiality
- Partial
- Category
- Cryptography & computation
Technical detail and cited results
Three routes lead to exact results.
-
Invariance. Kernels fix the reduction order for each output element regardless of batch size. Thinking Machines made RMSNorm, matrix multiplication and attention batch-invariant, the last with a fixed split size for the key-value dimension rather than a fixed number of splits S-1009. vLLM exposes this behind VLLM_BATCH_INVARIANT=1 on NVIDIA GPUs of compute capability 8.0 or higher and on Intel XPUs, in beta S-1013. SGLang integrated batch-invariant attention for its FlashInfer, FlashAttention 3 and Triton backends S-1012. LLM-42 instead decodes on a non-deterministic fast path and replays candidate tokens under a fixed-shape reduction schedule, rolling back any that are inconsistent S-1014.
-
Record and replay. Stock engines are deterministic but not invariant. Outputs are bitwise reproducible if the verifier knows the hardware model, the exact deployed weights, the parallelism topology (separately for prefill and decode), software versions including custom kernels, and the batch size of each forward pass S-0018 S-0020. Of these, only batch size changes during serving, and it costs one extra integer per forward pass to record S-0020. A software emulator reproduces the rounding of other GPU models by modelling tensor-core accumulation and kernel-specific reduction trees S-0020. Hawkeye reproduces tensor-core matrix multiplication exactly on a CPU for Ampere, Hopper and Ada Lovelace GPUs in FP16, BF16 and FP8 S-1010.
-
Reproducible operators. Gensyn's RepOps fixes the order of floating-point operations across hardware, so providers and a referee can reproduce the same result S-1809.
With exact replay, verification is pass/fail, and the chance of catching at least one false output in k samples is 1 − (1 − p)^k for a false-output rate p S-0020.
Claims and scope
A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.
The declared model is the one being served
Enables exact-match recomputation checks that the declared model, weights and software setup produced the outputs.
This compute runs inference, not training
Bit-exact recomputation of declared inference removes the tolerance an operator could hide other work in (S-0020).
Model weights have not left the facility
Removes the tolerance margin that steganographic exfiltration could use (S-0020).
Readiness for a stated use
Assessed use: reproducing open-model inference from receipts in Gensyn's information-market service
low confidence · current · assessed 2026-10-08 · rubric 1.1
This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.
Gensyn reports using exact replay in production so that anyone can check markets it settles with REE, but the production evidence comes from Gensyn, and no independent security evaluation of bit-exact verification has been published.
- R1 met: the verification claim, a covert-adversary threat model and the information a verifier needs are published S-0020.
- R2 met: public code predicts dense LLM blocks bit for bit on A100, L40, L40S and H100 GPUs running unmodified vLLM and Hugging Face engines, against a stated adversary S-0020. Batch-invariant modes are public in vLLM (beta) and SGLang S-1013 S-1012, and batch invariance was shown on a 235-billion-parameter model S-1009. Exact CPU reproduction of GPU matrix multiplication is peer-reviewed S-1010.
- R3 met through Gensyn's REE, on its developer's account. Gensyn reports that Delphi, its information-market app, is live on its mainnet, and that markets settled by open models inside REE produce receipts that anyone can re-run to verify the answer S-3022. Its service documentation also describes the REE judge receipts and mainnet deployment S-0075. REE is publicly available, and its README lists v0.8.0, released on 5 October 2026 S-1812. Gensyn does not label its releases alpha or beta S-3023. Eigen Labs' service does not count towards R3, because Eigen Labs launched it as a mainnet alpha S-3021. No party other than a developer is documented relying on exact replay for a verification decision. The evidence for replay of stock serving engines remains at R2 and covers the emulator's tested dense model blocks S-0020.
- R4 not met: as of September 2026 no independent audit, red-team or peer-reviewed security analysis of bit-exact verification has been published.
Confidence is low: R3 rests on Gensyn's account of its own service.
Evidence needed for the next level
An independent public evaluation (audit, red-team or peer-reviewed security analysis) of bit-exact verification that leaves no critical flaw open.
Limitations, flaws, and blockers
These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.
minor / open / open question
Some kernels remain genuinely nondeterministic
The bit-exact work separates kernels that are deterministic but not batch-invariant from truly nondeterministic ones that use atomic functions. Some integer de-quantization kernels use atomic additions and remain nondeterministic, so exact replay needs backends that avoid them S-0020.
significant / open / open question
Cross-hardware replay relies on reverse-engineered, closed behaviour
Emulating one GPU's rounding on another requires reverse-engineering tensor-core arithmetic and modelling proprietary kernel choices. Hawkeye covers a subset of NVIDIA architectures and states that attention and other higher-level operations need further reverse engineering S-1010. For the bit-exact emulator, a proprietary Hopper kernel family is an open edge case S-0020.
What still blocks use or stronger assurance
- S-1009S-1012
Batch-invariant kernels cost throughput: in Thinking Machines' Qwen3-8B test, an improved deterministic build took 42 s against 26 s for vLLM's default, and SGLang reports an average 34.35% slowdown on its FlashInfer and FlashAttention 3 backends.
- S-0020S-1013S-1814
Coverage is incomplete: the bit-exact emulator targets dense blocks on NVIDIA GPUs and excludes mixture-of-experts inference and training, and vLLM's batch-invariant mode is in beta, with open work on AMD hardware and speculative decoding.
- S-1008
Amodo's status page for the AI 2040 verification plan rates a reproducible inference stack for that plan as 'not started'.
- S-0020S-0018
Exact replay requires the prover to disclose weights, software versions, parallelism and batch sizes to whoever recomputes.
Connections in the research map
Complementary techniques
Implementations
Sources and provenance
- S-0020 / Tier B
Bit-Exact AI Inference Verification Without Performance Tradeoffs ↗
N. Cankaya · 2026 · ICML 2026 Workshop on Technical AI Governance Research
Supports: covert-adversary threat model; deterministic but non-invariant engines; required metadata; software emulator and its results; limitations; comparison with statistical schemes; best-paper award and public code (arXiv comments)
Locator: abstract; §1; results; limitations section; arXiv comments
Version and catalogue details - S-0015 / Tier B
Verifying LLM Inference to Detect Model Weight Exfiltration ↗
R. Rinberg, A. Karvonen, A. Hoover, D. Reuter, K. Warr · 2025 · arXiv
Supports: logging of inferences and random sampling for verification as components separate from recomputation
Locator: §4.2, assumptions 2 and 3 (v3)
Version and catalogue details - S-0018 / Tier B
A System Overview for Near-Term, Low-Trust AI Compute Verification ↗
N. Cankaya · 2026 · Machine Intelligence Research Institute
Supports: replay metadata in a low-trust verification system
Locator: §5.2.2
Version and catalogue details - S-0016 / Tier B
DiFR: Inference Verification Despite Nondeterminism ↗
A. Karvonen, D. Reuter, R. Rinberg, L. Marks, A. Garriga-Alonso, K. Warr · 2025 · ICML 2026 Workshop on Technical AI Governance Research
Supports: sources of benign nondeterminism; fixed-hardware determinism vs heterogeneous deployments; >98% token agreement
Locator: §2; §3; §7
Version and catalogue details - S-0017 / Tier C
Example Schemes for Verifying High-Stakes AI Agreements ↗
Amodo Design · 2026 · Amodo Design
Supports: floating-point non-associativity; fuzzy comparison in recomputation schemes
Locator: determinism discussion
Version and catalogue details - S-0067 / Tier C
Verification Plan ↗
R. Dean · 2026 · AI 2040
Supports: reproducibility required for packet correctness checks
Locator: Concrete inference-only retrofitting proposal
Version and catalogue details - S-1008 / Tier C
AI 2040 Plan A — Verification SITREP ↗
Amodo Design · 2026 · Amodo Design
Supports: status of reproducible inference stack
Locator: status items
Version and catalogue details - S-1009 / Tier C
Defeating Nondeterminism in LLM Inference ↗
H. He, Thinking Machines Lab · 2025 · Thinking Machines Lab: Connectionism
Supports: batch invariance as main cause; kernels made invariant; Qwen3-235B experiment; timings
Locator: whole post
Version and catalogue details - S-1010 / Tier A
Hawkeye: Reproducing GPU-Level Non-Determinism ↗
E. Badash, D. Boneh, I. Komargodski, M. Srivastava · 2026 · Proceedings of Machine Learning and Systems 8 (MLSys 2026)
Supports: exact CPU reproduction of tensor-core matrix multiplication; scope limits
Locator: abstract; §8; §9
Version and catalogue details - S-1011 / Tier B
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence ↗
DeepSeek-AI · 2026 · arXiv
Supports: provider-reported end-to-end batch-invariant and deterministic kernels
Locator: §3.3
Version and catalogue details - S-1012 / Tier C
Towards Deterministic Inference in SGLang and Reproducible RL Training ↗
The SGLang Team · 2025 · LMSYS Org blog
Supports: SGLang deterministic mode and its overhead
Locator: whole post
Version and catalogue details - S-1013 / Tier B
Batch Invariance (vLLM documentation) ↗
vLLM project · 2026 · vLLM documentation (GitHub, docs/features/batch_invariance.md)
Supports: vLLM batch-invariance flag, supported hardware (NVIDIA compute capability 8.0+, Intel XPUs), beta status
Locator: whole page
Version and catalogue details - S-1014 / Tier B
LLM-42: Enabling Determinism in LLM Inference with Verified Speculation ↗
R. Gond, A. K. Kamath, R. Ramjee, A. Panwar · 2026 · arXiv
Supports: scheduling-based determinism alternative that mostly reuses existing kernels
Locator: abstract
Version and catalogue details - S-1813 / Tier B
thinking-machines-lab/batch_invariant_ops (GitHub repository) ↗
Thinking Machines Lab · 2025 · GitHub
Supports: Thinking Machines' batch-invariant kernel library (MIT)
Locator: README
Version and catalogue details - S-1814 / Tier C
[Feature]: Batch Invariant Feature and Performance Optimization (vLLM issue #27433) ↗
vLLM project contributors · 2025 · GitHub (vllm-project/vllm issues)
Supports: vLLM developers' statement that batch-invariance support is based on the Thinking Machines post; open work on AMD hardware, NVFP4 and speculative decoding
Version and catalogue details - S-1809 / Tier B
Verde: Verification via Refereed Delegation for Machine Learning Programs ↗
A. Arun, A. St. Arnaud, A. Titov, B. Wilcox, V. Kolobaric, M. Brinkmann, O. Ersoy, B. Fielding, J. Bonneau · 2025 · arXiv
Supports: Verde refereed delegation and RepOps reproducible operators
Locator: abstract; §3.2
Version and catalogue details - S-1810 / Tier C
Verde Verification System In Production ↗
O. Ersoy · 2025 · Gensyn research blog
Supports: Gensyn's report of Verde and RepOps in production (provider-reported)
Version and catalogue details - S-1812 / Tier B
gensyn-ai/ree: Gensyn Reproducible Execution Environment (GitHub repository) ↗
Gensyn · 2026 · GitHub
Supports: REE public as binaries with an MIT-licensed SDK; v0.8.0 released on 5 October 2026 (provider-reported)
Locator: README; patch notes
Version and catalogue details - S-3020 / Tier B
EigenAI: Deterministic Inference, Verifiable Results ↗
D. Ribeiro Alves, V. Patankar, M. Pereira, J. Stephens, N. Vaziri, S. Kannan · 2026 · arXiv
Supports: EigenAI deterministic engine and optimistic re-execution protocol; same-SKU and A100-vs-H100 determinism results (provider-reported)
Locator: abstract; Table 5
Version and catalogue details - S-3021 / Tier C
EigenCloud Brings Verifiable AI to Mass Market with EigenAI and EigenCompute Launches ↗
EigenCloud · 2025 · Eigen Labs blog
Supports: EigenAI mainnet alpha launch; stake not yet exposed to slashing (provider-reported)
Version and catalogue details - S-3022 / Tier C
Building Delphi: Pricing, Settlement, and Agentic Trading ↗
D. Jedamski · 2026 · Gensyn blog
Supports: Delphi live on Gensyn's mainnet; REE settlement receipts that anyone can re-run (provider-reported)
Locator: settlement section
Version and catalogue details - S-3023 / Tier B
Reproducible Execution Environment (REE) (Gensyn documentation) ↗
Gensyn · 2026 · Gensyn documentation
Supports: REE releases carry no alpha or beta label (provider-reported)
Locator: whole page
Version and catalogue details - S-0075 / Tier B
What is Delphi? (Delphi documentation) ↗
Gensyn · 2026 · Delphi documentation
Supports: Delphi mainnet service and REE judge receipts that anyone can re-run (provider-reported)
Locator: Settled by AI, Verifiable by Anyone; Where Delphi Runs
Version and catalogue details - S-3500 / Tier A
TAO: Tolerance-Aware Optimistic Verification for Floating-Point Neural Networks ↗
J. Yao, H. Su, T. Liao, Z. Cheng, H. Zhang, X. Wang, P. Viswanath · 2026 · Proceedings of the 21st European Conference on Computer Systems (EuroSys 2026), pp. 1515-1532
Supports: TAO: operator-level tolerance bounds instead of bitwise equality on heterogeneous hardware; Merkle-anchored dispute game; Ethereum testnet deployment
Locator: abstract; §1; evaluation
Version and catalogue details
- Source review date
- 2026-09-25
- Drafted by (source map)
- ai
- Review handles (source map)
- codex-review