I-0016 / Cryptography & computation

Batch-invariant inference kernels (Thinking Machines)

Open-source kernels from Thinking Machines Lab that make LLM outputs independent of batch size, adopted in vLLM and SGLang to give reproducible inference.

R2 DemonstratedSource reviewed 2026-09-25Provider-reported evidence

01 / The mechanism and its boundary

What the technique establishes

Batch-invariant kernels make a language model give the same output for the same input whatever else the server is processing. Thinking Machines Lab identified changing batch sizes as the main reason LLM endpoints are nondeterministic, and published kernels whose reduction order does not depend on the batch. With them, 1,000 temperature-zero completions from a 235-billion-parameter model were identical. SGLang built its deterministic mode on the kernels, and vLLM's batch-invariant mode, still in beta, is based on the same work. For verification, identical outputs let a verifier that runs the same model, engine and hardware require an exact match instead of a tolerance. The kernels were built for reproducibility, not verification, and no verification system is documented using them. They slow inference: by 61.5% in Thinking Machines' test, whose vLLM integration was not yet heavily optimised, and by 34% on average in SGLang.

Threat model
Cooperative prover
Adversarial evaluation
None
Hardware needed
None
Prover cooperation
Required
Confidentiality
Revealing
Category
Cryptography & computation

Claims and scope

A direct link identifies the intended claim. A supporting link supplies part of the evidence. Neither establishes that a complete verification system has been demonstrated.

Readiness for a stated use

R2 Demonstrated

Assessed use: exact recomputation of served outputs by a verifier, with a cooperating provider

low confidence · current · assessed 2026-09-25 · rubric 1.1

This is the source map’s editorial assessment. Production use is not evidence of resistance to every adversary.

Public kernels made a 235-billion-parameter model's outputs identical across 1,000 runs and underpin the deterministic modes of two production engines, but no verification system is documented using them.

  • R1 met: Thinking Machines sets out the design and the property it gives, outputs that do not depend on batch size S-1009. The condition for identical outputs, a fixed model, inference implementation and device, is stated in the verification literature S-0016.
  • R2 met: the code is public under an MIT licence, with a deterministic vLLM example S-1813. With the kernels, 1,000 temperature-zero completions from Qwen3-235B were identical, against 80 unique completions without them S-1009.
  • R3 not met for this use. The kernels have reached production engines: SGLang builds its deterministic mode on them S-1012, and vLLM's developers state that its batch-invariant mode, labelled beta, is based on the same work S-1814 S-1013. The kernels were built for reproducibility, such as on-policy reinforcement learning S-1009. Gensyn's REE and Eigen Labs' EigenAI, which build verification services on exact replay, use their own kernels S-1809 S-3020. No party is documented relying on these kernels for a verification decision.
  • R4 not met: no independent security evaluation has been published.

Confidence is low, because whether an engine's deterministic mode counts as production-grade for verification is a judgment call.

Evidence needed for the next level

  • A production-grade verification stack that uses these kernels for exact-match checks, or a party other than the developers relying on such checks for a verification decision.

Limitations, flaws, and blockers

These are attributed assessments from the source map. Absence of a listed flaw is not a security guarantee.

No flaws are recorded in this snapshot. See the scope, assumptions, and blockers below.

What still blocks use or stronger assurance

  1. Batch invariance costs throughput: on Qwen3-8B the improved deterministic build took 42 s against 26 s for vLLM's default, and SGLang reports an average slowdown of 34.35% on its FlashInfer and FlashAttention 3 backends.

    S-1009S-1012
  2. Outputs are identical only while the model, inference implementation and device stay fixed, so provider and verifier must run the same stack.

    S-0016

Connections in the research map

Mechanisms implemented

Concepts used

Organizations and developers

Sources and provenance

  1. S-1009 / Tier C

    Defeating Nondeterminism in LLM Inference ↗

    H. He, Thinking Machines Lab · 2025 · Thinking Machines Lab: Connectionism

    Supports: main cause of nondeterminism; operations made batch-invariant; fixed split size for attention; vLLM integration; Qwen3-235B and Qwen3-8B results; on-policy RL motivation (provider-reported)

    Locator: whole post

    Version and catalogue details
  2. S-1813 / Tier B

    thinking-machines-lab/batch_invariant_ops (GitHub repository) ↗

    Thinking Machines Lab · 2025 · GitHub

    Supports: MIT-licensed library, replaced PyTorch operations, deterministic vLLM example

    Locator: README

    Version and catalogue details
  3. S-1012 / Tier C

    Towards Deterministic Inference in SGLang and Reproducible RL Training ↗

    The SGLang Team · 2025 · LMSYS Org blog

    Supports: SGLang's integration of the kernels, its attention kernels, the 61.5% figure for the original post and SGLang's average slowdown

    Locator: whole post

    Version and catalogue details
  4. S-1013 / Tier B

    Batch Invariance (vLLM documentation) ↗

    vLLM project · 2026 · vLLM documentation (GitHub, docs/features/batch_invariance.md)

    Supports: vLLM batch-invariant mode, beta status, supported hardware (NVIDIA compute capability 8.0+, Intel XPUs), tested model families

    Locator: whole page

    Version and catalogue details
  5. S-1814 / Tier C

    [Feature]: Batch Invariant Feature and Performance Optimization (vLLM issue #27433) ↗

    vLLM project contributors · 2025 · GitHub (vllm-project/vllm issues)

    Supports: vLLM developers' statement that batch-invariance support is based on the Thinking Machines post; open work items

    Version and catalogue details
  6. S-0016 / Tier B

    DiFR: Inference Verification Despite Nondeterminism ↗

    A. Karvonen, D. Reuter, R. Rinberg, L. Marks, A. Garriga-Alonso, K. Warr · 2025 · ICML 2026 Workshop on Technical AI Governance Research

    Supports: identical results only with fixed model, implementation and device; statistical verification for heterogeneous stacks

    Version and catalogue details
  7. S-1809 / Tier B

    Verde: Verification via Refereed Delegation for Machine Learning Programs ↗

    A. Arun, A. St. Arnaud, A. Titov, B. Wilcox, V. Kolobaric, M. Brinkmann, O. Ersoy, B. Fielding, J. Bonneau · 2025 · arXiv

    Supports: RepOps, Gensyn's own reproducible operators (rationale comparison)

    Locator: §3.2

    Version and catalogue details
  8. S-3020 / Tier B

    EigenAI: Deterministic Inference, Verifiable Results ↗

    D. Ribeiro Alves, V. Patankar, M. Pereira, J. Stephens, N. Vaziri, S. Kannan · 2026 · arXiv

    Supports: EigenAI's own deterministic kernels (rationale comparison; provider-reported)

    Locator: abstract; kernels section

    Version and catalogue details
Source review date
2026-09-25
Drafted by (source map)
ai
Review handles (source map)
codex-review