Four things happen between your prompt and each word of the reply (technically a token, but let's keep it simple). The forward pass scores every token in the vocabulary. Temperature reshapes those scores into a probability distribution. Top-k and top-p decide which candidates stay eligible. The draw picks one, and the draw is the only randomness in that account, with the seed as the thing that governs it.

Here is what you would expect. Turn the temperature down to zero and the distribution collapses onto a single token, so the draw is still there but has nothing left to decide. Or leave the temperature alone and pin the seed, and the draw makes the same choice every time. The same prompt should come back with the same answer every time.

It does not, and the gap is not small. This article documents what happened when I checked.

Note: The companion piece to this article is What Changes When Part of Your System Is Non-Deterministic. To make full sense of what follows, read it first. It argues that non-deterministic is a weak word that hides much of what you need to know. It also explains temperature, top-k, top-p and the seed, the four hyper-parameters this article refers to.

So I measured it

I sent the same prompt one hundred times to gpt-4.1-mini, pinned to its 2025-04-14snapshot, and then did it five more times with the sampling controls changed: two temperatures, crossed with no seed, one fixed seed, and a different seed on every call. Six hundred calls, made one after another rather than in parallel, because firing them concurrently could have put my own requests in the same server batch, and batch composition was one thing I wanted to leave alone.

The prompt was the same every time, with no system message:

List three reasons a flaky test might pass on a retry. One sentence each.

Two replies count as different when the strings are not identical, character for character, with no normalisation applied.

As a separate check, I lower-cased every reply, stripped the punctuation and collapsed the whitespace, and none of the six counts below changed. Every difference is a difference in words.

The following are the readings from the experiment:

Setting

Different replies, out of 100

temperature 0, no seed

27

temperature 0, seed=42 on every call

30

temperature 0, a different seed each call

27

temperature 1.0, no seed

100

temperature 1.0, seed=42 on every call

26

temperature 1.0, a different seed each call

100

Figure 1

Each row in that chart is the hundred replies for one setting, sorted so that identical replies sit together. One block is one distinct reply, and its width is how many times that reply came back. So the number of blocks in a row is the number in the margin, and you can count them.

Read the temperature 1.0 rows first, because the seed behaves there the way you would expect. With no seed, and with a different seed each time, the rows are a hundred hairlines: not one reply came back twice. Hold the seed fixed instead and the row collapses into a few wide blocks, the largest of them thirteen calls returning the same answer, word for word. The seed is honoured, and it is doing real work.

It is also not doing the job you hired it for. A seed that governed the output would return one reply a hundred times. It returned twenty-six. The seed fixes the path through the distribution it is handed, and the likeliest reading is that it is not handed the same distribution twice.

Now read the temperature 0 rows, where the same three settings produce 27, 30 and 27. The seed has stopped mattering. That is what the pipeline predicts: with the distribution collapsed the draw is down to a single candidate, so there is nothing left for a seed to govern. It is satisfying to see a prediction survive contact with a hundred calls, and this one did.

And then the thing that should not be there. Four of these six settings land between 26 and 30 different replies. Those four reach it by one of two routes: at temperature zero the draw has a single candidate and makes no choice, and with a fixed seed at temperature 1.0 it makes the same choice every time. Different routes, and both stop in the same place: around 27 different replies out of a hundred.

That floor is the variation left when the sampler has nothing more to contribute. Whatever is left, no sampler control reaches it. So where is it?

Where the replies part company

Figure 2

The above chart is the same data read along the reply instead of across it. For each character position it asks how many of the hundred replies are still word-for-word identical to one another, counting the largest group that still agrees. Where each line finally flattens is the width of the widest block in the earlier chart, so the two meet at the right-hand edge. For example, follow the solid line for temperature 0, no seed, to the right edge and it settles on 22. That's the same 22 as the widest block in the earlier chart, the largest number of replies that were identical for this setting.

At temperature zero with no seed, ninety-nine of the hundred replies are identical for the first eighty-four characters. Then a step down to seventy-one, another to thirty-six, and twenty-two of them are identical the whole way through. Each of those steps is one place where the top-scoring token changed, and everything downstream of a step is a different sentence, or several.

The two runs at temperature 1.0 without a fixed seed sit on top of each other and fall off a cliff. By character eighty-four, where temperature zero with no seed still has ninety-nine replies in agreement, they have nine.

The steps are the point. The three temperature-zero lines do not drift; each agrees exactly, then flips, then agrees exactly again on the other side of the flip. That variation is not something the decoding controls can reach, because at temperature zero they have nothing left to decide.

My experiment only establishes this much, and the only part that I can defend is this:

The settings that we control in a request have an impact on the variation, but the variation does not go away. It's still significant and can't be removed with what our request can control.

What I cannot show you

What this experiment does not establish is the cause, and this is where a measurement quietly turns into a story if nobody stops it.

The draw has nothing left to contribute, so the standard account looks at the step before it. The forward pass scores every token in the vocabulary, and it does that with arithmetic on the same numbers every time. That should give the same answer twice. The standard account says it does not, because floating-point addition is not associative. Add the same numbers in a different order and the last bits can come out different. That order depends on how the work was batched, and the batch depends on who else was calling the API at that moment. I find that account convincing, and it is argued in detail by people who build inference servers. But I did not measure it.

My runs also crossed more than one backend configuration. The system_fingerprint on the responses took four distinct values across the six settings, which is a second possible cause I cannot separate from the first. From outside the serving stack I can show you that a floor exists and that no sampler control reaches it. I cannot open the box and show you why. Also, these are the readings for the one prompt I ran. Six hundred calls make the floor of this prompt—my prompt—visible and say nothing about where your prompt is going to settle.

I could not diagnose it because I do not own the infrastructure. Neither do you. Non-determinism is relative to the boundary you can see. The provider can see and control batch composition, because it is an input on their side. You cannot see it, so the same variation that is ordinary bookkeeping on their side is irreducible on yours.

Does it matter?

Running an experiment and reading its numbers are two different jobs. Deciding whether those numbers describe a problem is a third. The readings are the same for everyone. What they cost you depends on what you were going to do with the reply.

The differences found in my experiment were not dramatic in themselves. For my sample prompt that asks for three reasons a flaky test might pass on a retry, two replies at temperature 0 with no seed, each of which came back once, are word-for-word identical for 347 characters. After that, they had their differences over how the sentence ends: "is reset before the retry" or "resets between retries". Both say nearly the same thing.

A semantic check would call them near-paraphrases and be right to. An exact-match assertion fails on the pair every time, objecting to wording it should ignore. The semantic check is right about this pair, and I do not trust it to object to a reply that is wrong and fluent. Neither of those is right or wrong. 

An oracle is a choice about what we are willing to call a difference, and that choice belongs to what we are testing rather than to the oracle.

Maybe we need both, and more than both. Maybe we keep reaching for matching and mathematics because that is what we have. And maybe, a response from a language model does not lend itself to them all the time.

Can you make it deterministic?

Not from where you are sitting, although a good deal of published advice implies you can. Set the temperature to zero, pin the seed, and pin the model snapshot. That recipe still returned thirty different replies in a hundred.

I used the two controls people reach for in practice. This model's API exposes another hyper-parameter that touches sampling, top-p, and it cannot help. Top-p keeps the highest-scoring candidates and discards the tail. It can remove a token the draw might otherwise have picked, but it always keeps the highest-scoring one, and at temperature zero that is the token being chosen.

I ran it rather than assume it. Three hundred more calls at temperature zero, with top-p at 0.05, at 0.5, and left at its default of 1.0, returned 30, 31 and 32 different replies. No seed was pinned for these, because the six settings above had already shown the seed doing nothing at temperature zero. Truncating to the top five per cent of the probability mass and not truncating at all land in the same place.

Temperature, top-p and the seed all act after the forward pass, and none of them reaches whatever is varying.

It would be easy to swing the other way and decide that a language model is simply unpredictable. That is wrong too. The variation is not coming only from the model, and it is not coming from anything you can set on the request. The account I find convincing puts it in the serving stack, which batches your request in with other people's and reorders the arithmetic. That is the part I did not measure. Somebody who controls that stack can take it out. Thinking Machines published batch-invariant kernels that make vLLM return bitwise identical output run after run, which settles the question of whether it is possible.

So the expectation is worth carrying: the final leg of determinism is a property of the infrastructure, not a parameter on the request. If you own the serving stack you can buy it, and you will pay for it in throughput. If you are calling somebody else's API, no combination of the settings on offer will change it. What would have to change is on their side of the boundary, and nothing I tried from this side reached it.

Should determinism be a universal ask?

Pinning the output has a real use, and it is worth saying so before I argue with it.

If you want to know that something changed between two runs, comparing strings will tell you about it, and this is the cheapest method. The mistake is treating a match as a guarantee rather than a signal to go and look. A test strategy built on pinning the output is not being careful. It is assuming a guarantee nobody gave you. It will hold until it doesn't, and you will not get a warning.

Before spending more effort on this: what would you do with a deterministic model if you had one? The same string every call is reproducibility, not correctness. A test that asserts one exact answer passes just as happily when the model is reliably wrong. For the kind of problems language models are meant to solve, what you want is probably a way to tell a good answer from a bad one. Where a right answer can be worded more than one way, neither exact matching nor a semantic check gives you that on its own. This line of thinking is especially critical for automated assertions, because a human oracle cannot scale to the volume of testing a language-model-augmented solution truly needs.

Where I stopped

This article is the honest reading of the range of experiments I did in the process of writing it. I got as far as the sampler and stopped there. Nothing I could reach held the output still, and what would have to change is not exposed on the request.

For some systems the journey ends because there is nothing further to learn. For a hosted model it ends sooner, and not because the remaining distance is hard. It ends because the part that would have to change is not ours.

That leaves the question the next article opens on. If you cannot pin the answer, what do you assert instead? That one is What to Assert When You Cannot Assert the Answer.

Acknowledgements

With thanks to Claude, and to the GPT, Gemini, Grok, DeepSeek, GLM, Kimi and Qwen model families, who between them made this harder to publish and better to read. How the research and the writing were done is set out in full in the companion article, What Changes When Part of Your System Is Non-Deterministic.

The positions here are mine, and so is anything still wrong with them.