Run it twice. You may get two answers.
That sentence should be uncontroversial for anyone who has read the API documentation carefully, and it comes as a genuine surprise to most people who use these systems in professional work. The assumption underneath a great deal of AI practice, that a fixed prompt against a fixed model at temperature zero is a repeatable operation, is not true of any major hosted service. The vendors say so in their own reference material. The consequences run further than most governance frameworks have followed.
Why does temperature zero not mean deterministic?
Because temperature governs sampling, and sampling is not the only source of variation.
Anthropic's Create a Message API reference is direct about it: "Note that even with temperature of 0.0, the results will not be fully deterministic." There is no seed parameter to reach for, and the same page now marks temperature itself deprecated: models released after Claude Opus 4.6 "do not support setting temperature" and reject any value other than 1.0 with a 400 error. On Anthropic's newest models, temperature zero is not a dial you can turn.
OpenAI's advanced usage guide states that "Chat Completions are non-deterministic by default (which means model outputs may differ from request to request)", offers a seed for "mostly" deterministic results, and adds that "determinism may be impacted due to necessary changes OpenAI makes to model configurations on our end", surfaced through the system fingerprint field.
Google's content generation parameters documentation says that even when a seed is fixed to a specific value the model "makes a best effort to provide the same response for repeated requests" and that "Deterministic output isn't guaranteed", and that at temperature zero responses are "mostly deterministic, but a small amount of variation is still possible".
Three vendors, three separate disclaimers, one shared position. This is documented behaviour rather than a defect.
The scale of it is easier to grasp with a measurement. Thinking Machines Lab sampled 1,000 completions at temperature zero from one open model using an identical prompt, generating 1,000 tokens each. The result: "we generate 80 unique completions, with the most common of these occuring 78 times". More usefully for anyone thinking about spot checks, the completions "are actually identical for the first 102 tokens". Divergence appears well into the response, which is exactly where a human reviewer has stopped reading closely.

What actually causes the variance?
Not randomness in the model. Other people's traffic.
The Thinking Machines Lab analysis identifies the mechanism plainly: "the primary reason nearly all LLM inference endpoints are nondeterministic is that the load (and thus batch-size) nondeterministically varies". Requests arriving at a serving endpoint are grouped into batches for efficiency, and the size of that batch depends on how busy the endpoint is at that instant.
Batch size then propagates into the arithmetic. SGLang's documentation states the chain compactly: "Different batch sizes cause GPU kernels to split reduction operations differently, leading to different addition orders. Due to floating-point non-associativity... this produces different results even for identical inputs." Adding the same set of numbers in a different order produces a very slightly different sum, that difference shifts token probabilities by a fraction, and occasionally that fraction is enough to change which token wins.
Which leads to the observation that should reframe how governance teams think about this. In the lab's words, "our request's output does depend on the parallel user requests", and from the user's point of view "the other concurrent users are not an 'input' to the system but rather a nondeterministic property of the system".
An organisation running a regulated process on a hosted model is running a process whose output is influenced, in a small but real way, by unrelated strangers hitting the same endpoint in the same second. That is not a sentence most AI risk registers contain.

Can it be fixed?
On infrastructure you control, increasingly yes. On the API you buy, no.
Both major open source serving stacks now ship deterministic execution as a documented feature. vLLM's batch invariance documentation describes it as ensuring "that the output of a model is deterministic and independent of the batch size or the order of requests in a batch", enabled by an environment variable, requiring reasonably recent NVIDIA hardware or an Intel XPU on the Triton backend. The work landed in named releases, v0.11.1 on 18 November 2025 and v0.12.0 on 3 December 2025. SGLang exposes the same capability behind an explicit deterministic inference flag, restricted to three attention backends.
The cost is stated honestly by both. vLLM notes: "Enabling batch invariance may impact performance compared to the default non-deterministic mode. This trade-off is intentional to guarantee reproducibility." In the lab's own benchmark the same workload took 26 seconds under default vLLM, 55 seconds with unoptimised deterministic kernels, and 42 seconds after an improved attention kernel. Determinism is bought with throughput and with features. vLLM's implementation notes say it disables certain optimisations that can introduce non-determinism, naming custom all-reduce operations in tensor parallel mode, and one of SGLang's three permitted backends gives up the radix cache to run deterministically at all.
Two qualifications matter before anyone treats this as solved. The vLLM feature is documented as beta and validated against a named list of models, so it is not yet a production guarantee. And more importantly, none of this reaches the hosted endpoints most Australian organisations actually consume. If you buy an API, you inherit the vendor's batching, and the vendor has told you in writing that determinism is not on offer.

How much does this actually vary in practice?
It depends heavily on the task, and the honest answer is that the public evidence is thin.
A workshop paper from Raffi Khatchadourian and Rolando Franco, accepted to the AI4F workshop at ACM ICAIF 2025, tested five models across three regulated financial tasks over 480 runs and reported that "smaller models (Granite-3-8B, Qwen2.5-7B) achieve 100% output consistency at T=0.0, while GPT-OSS-120B exhibits only 12.5% consistency (95% CI: 3.5-36.0%)". That confidence interval is wide and the sample behind it is small, sixteen runs per condition, so it should be read as an indication rather than a measurement. The more durable finding in the same work is the shape of the exposure: structured generation stayed stable, while retrieval augmented tasks drifted considerably.
That pattern is intuitive once the mechanism is clear. The more a task depends on a long chain of generated reasoning, the more opportunities a single flipped token has to send the response somewhere else.
What does this break?
Three things, and the first is the most widely practised.
Re-running a prompt is not verification. The standard check, when an output looks wrong, is to run it again and see whether the same thing comes back. Under non-determinism that procedure cannot distinguish between a model that was wrong and a model that was unlucky, and a second run that looks fine provides no evidence about the first. Worse, the divergence tends to appear deep in the response, after the point where a reviewer comparing two outputs has decided they match.
A single evaluation score is one draw, not a measurement. If a model is benchmarked once on an internal eval set, the number produced is a sample from a distribution whose width nobody has measured. Comparing two models on one run each is comparing two draws. This does not invalidate private evaluation, which remains far better than trusting a public leaderboard, but it does mean the run count is part of the method.
Schema validity is not substantive agreement. Constrained decoding against a JSON schema guarantees the shape of an answer. Two responses can both parse cleanly and still differ on the field that carries the decision.
Does this compound across agent steps?
Yes, and this part has now been measured rather than merely reasoned about.
An agent run is a sequence of model calls where each call conditions on the output of the last. If any single call can diverge, then a divergence early in the sequence does not simply alter one sentence, it changes the input to every step that follows. A tool gets selected differently, a different document gets retrieved, and the two runs stop being comparable well before the final answer.
Khatchadourian returned to the question with a harness built for exactly this, published as the Determinism-Faithfulness Assurance Harness. It ran "4,700+ agentic runs (7 models, 4 providers, 3 financial benchmarks with 50 cases each at T=0.0)" and separates two things most teams collapse into one: whether the agent reached the same decision, and whether it got there by the same route. On the compliance triage benchmark, signature determinism, meaning the fraction of cases where every run produced an identical tool sequence and arguments, was 40.9 percent for Gemini 2.5 Pro and 44.0 percent for both Claude Sonnet 4 and Claude Opus 4.5, against decision determinism of 68.2, 82.0 and 72.0 percent for the same three. The paper's own reading is that "tool-path variance" rather than decision variance is "the primary source of non-reproducibility in agentic settings".
Two things follow for regulated work. Decision determinism and accuracy were "not detectably correlated" in that sample, so a model that reliably returns the same answer is not thereby a model that returns the right one. And an agent can arrive at the same decision twice by two different routes, which means the run log matters more than the prompt, and matching final answers is not evidence that the same process produced them.
Where does this land for Australian work?
On record keeping, and the framing is already available.
The Voluntary AI Safety Standard was the usual reference point here, and its Guardrail 9 still asks organisations to "keep and maintain records to allow third parties to assess compliance with guardrails". The Department of Industry, Science and Resources has since moved past it. Its own page for the standard now carries a notice that the Guidance for AI Adoption, first published 21 October 2025 and since revised, is "updated and simplified guidance for industry" that "evolves the Voluntary AI Safety Standard". Neither document is law and neither creates a new legal obligation, but the current guidance states the same expectation more plainly: "Keep clear records of governance decisions, testing, incidents and monitoring." Its fifth essential practice, test and monitor, asks organisations to "clearly document tests and outcomes to support external audits and oversight".
Read those alongside non-determinism and a specific operational rule falls out. If the output cannot be regenerated, then the record has to be captured at the moment the decision is made. An assurance process that plans to reconstruct what the model said by re-running the prompt later is planning to produce a different document and call it evidence.
The timing is worth noting too. From 10 December 2026, an APP entity that has arranged for a computer program to use personal information to make a decision that could reasonably be expected to significantly affect an individual's rights or interests will need to describe that use in its privacy policy. Describing a system is not the same as evidencing an individual decision, but an organisation that cannot produce the second will find the first uncomfortable to write.
Bottom line
Reproducibility was never promised. Every major vendor documents that identical requests can produce different answers, the cause is other people's traffic rather than anything in your control, and the fixes that exist apply to serving stacks you run yourself. The practical response is not to chase determinism. It is to stop relying on procedures that quietly assume it: capture the artefact at decision time, evaluate over repeated runs, and treat a second run as a second opinion rather than a confirmation.
Do this Monday
- Find every process where the verification step is described as re-running the prompt, and replace it with a check against the source material instead
- Turn on artefact capture at the point of decision for any AI-assisted output that could be questioned later, storing the response rather than the means of regenerating it
- Re-run your internal evaluation set at least five times per model and record the spread, not just the score
- Ask your AI vendors in writing what they guarantee about output stability, and file the answer next to the contract
- Where a workflow genuinely requires reproducibility, check whether self-hosted serving with batch invariant execution is viable, and price the throughput cost before proposing it
References
- Anthropic, Create a Message, Messages API reference, accessed 21 September 2026. https://platform.claude.com/docs/en/api/http/messages/create
- OpenAI, Advanced usage, API guides, accessed 21 September 2026. https://developers.openai.com/api/docs/guides/advanced-usage
- Google Cloud, Content generation parameters, accessed 21 September 2026. https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/capabilities/content-generation-parameters
- Thinking Machines Lab, Defeating Nondeterminism in LLM Inference, 10 September 2025. https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/
- vLLM, Batch Invariance, project documentation, accessed 21 September 2026. https://docs.vllm.ai/en/latest/features/batch_invariance/
- SGLang, Deterministic Inference, project documentation, accessed 21 September 2026. https://docs.sglang.io/advanced_features/deterministic_inference.html
- Raffi Khatchadourian and Rolando Franco, LLM Output Drift: Cross-Provider Validation and Mitigation for Financial Workflows, AI4F workshop at ACM ICAIF '25, Singapore, 15 to 18 November 2025, arXiv 2511.07585. https://arxiv.org/abs/2511.07585
- Raffi Khatchadourian, Replayable Financial Agents: A Determinism-Faithfulness Assurance Harness for Tool-Using LLM Agents, arXiv 2601.15322 version 2, 7 March 2026, to appear in the 2nd ICLR Workshop on Advances in Financial AI. https://arxiv.org/abs/2601.15322
- National Artificial Intelligence Centre, Voluntary AI Safety Standard: the 10 guardrails, published 5 September 2024, page updated 1 December 2025. https://www.industry.gov.au/publications/voluntary-ai-safety-standard/10-guardrails
- National Artificial Intelligence Centre, Guidance for AI adoption: implementation guidance, current version published 5 May 2026, accessed 21 September 2026. https://www.ai.gov.au/staying-safe-and-responsible/essential-ai-practices/guidance-ai-adoption-implementation-guidance
- OAIC, Australian Privacy Principles guidelines, Chapter 1: APP 1, updated 3 October 2025. https://www.oaic.gov.au/privacy/australian-privacy-principles/australian-privacy-principles-guidelines/chapter-1-app-1-open-and-transparent-management-of-personal-information
TheAICommand. Intelligence, At Your Command.



