Close Menu
AI News TodayAI News Today

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    One with the world? A new look at brains transformed by psychedelics.

    AI Is Getting Really Good at Messing With Cybercriminals

    Danu Robotics’ fight to build a better recycling robot

    Facebook X (Twitter) Instagram
    • About Us
    • Contact Us
    Facebook X (Twitter) Instagram Pinterest Vimeo
    AI News TodayAI News Today
    • Home
    • AI News
    • AI Reviews
    • AI Tools
    • AI Tutorials
    • Chatbots
    • Free AI Tools
    • Artificial Intelligence
    AI News TodayAI News Today
    Home»AI Tools»Why Temperature 0 Isn’t Deterministic
    AI Tools

    Why Temperature 0 Isn’t Deterministic

    By No Comments10 Mins Read
    Share Facebook Twitter Pinterest LinkedIn Tumblr Reddit Telegram Email
    Why Temperature 0 Isn't Deterministic
    Share
    Facebook Twitter LinkedIn Pinterest Email

    In September 2025, Horace He and colleagues at Thinking Machines Lab ran a simple experiment. They sent the prompt “Tell me about Richard Feynman” to Qwen3-235B 1,000 times at temperature 0 and asked for 1,000 tokens each time. Temperature 0 means the model always picks its most probable token, so you would expect 1,000 identical answers. They got 80 unique completions.

    What caught my eye is where the answers split. All 1,000 completions were identical for the first 102 tokens. At token 103, 992 of them continued with “Queens, New York” and 8 with “New York City.” Eight runs in a thousand flipped at a single token, after 102 tokens where none did.

    Their post explains why outputs differ at all. This article asks a different question: how often should a token flip, and why does that risk arrive in sudden spikes? I’ll derive a one-line formula, check it with a simulation, and test it on a real model. Everything runs on a laptop, and the code is in the article.

    Where the noise comes from

    At temperature 0 the model picks the token with the highest logit (its raw score). That is deterministic mathematics, but your computer computes logits in floating point, where addition is not associative:

    >>> 0.1 + 0.2 + 0.30.6000000000000001>>> 0.1 + (0.2 + 0.3)0.6

    At scale it gets worse. I summed 100,000 random float32 numbers three ways. Forward order gave −90.82497, reverse order gave −90.82674, and NumPy’s pairwise sum gave −90.82513 (a float64 reference gives −90.82508). Same numbers, three answers.

    A neural network is billions of such sums, so the order the hardware uses matters. A popular explanation blames GPU threads that finish in random order. He et al. show that this isn’t the main cause: a typical LLM forward pass contains essentially no atomic adds, and the same kernel on the same input returns the same bits every time. The real culprit is that many kernels aren’t batch-invariant. Their reduction order changes with the batch size, and on a shared server the batch size depends on how many other people are sending requests at that moment. Your answer depends on strangers.

    The authors note that this isn’t specific to GPUs, and that with batch-invariant kernels all 1,000 Feynman completions came out identical. The fix has a price: in their serving test (Qwen-3-8B on one GPU), the unoptimized deterministic version took 55 seconds against 26 for default vLLM, and 42 after an improved attention kernel.

    Only near-ties matter

    Here is the part that interests me as a mathematician. Let z₁ and z₂ be the two largest logits, and call their difference the gap, M=z1−z2≥0M = z₁ − z₂ ≥ 0M=z1​−z2​≥0. Numerical noise changes that gap by a small error Δ, and the argmax flips only if M+Δ<0M + Δ < 0M+Δ<0.

    A token with a gap of 8 logits will never flip, however noisy the kernel. Only near-ties are at risk. Averaging over all positions, with f the density of gaps, the flip probability is

    p = ∫ P(Δ < −m) f(m) dm

    If the noise is small, f is roughly constant, equal to f(0), over the narrow range where P(Δ<−m)P(Δ < −m) P(Δ<−m) isn’t negligible. The integral of P(Δ<−m)P(Δ < −m)P(Δ<−m) over m≥0m ≥ 0m≥0 is the expected negative part of the error, which is E∣Δ∣/2E|Δ|/2E∣Δ∣/2 for symmetric noise. So

    p ≈ f(0) · E|Δ| / 2

    The flip probability per token is how crowded the neighbourhood of a tie is, times how large the noise is. The first factor belongs to the model and the text. The second belongs to your hardware, precision and kernels. This is an elementary small-noise argument, so I make no claim of novelty. A related formalization, the “background temperature” of Messina and Scotta (TMLR, 2026), treats such perturbations as an effective temperature.

    Schematic. The orange region holds the positions whose gap is smaller than a given error. Because the error has a random sign and size, averaging gives f(0)·E|Δ|/2.

    Does the formula hold?

    Before touching a real model, I checked the formula on simulated gaps with a known density and Gaussian noise of known size σ σσ. The core is nine lines:

    def flip_rate(sampler, sigma, n, rng, chunk=5_000_000):    flips, done = 0, 0    while done < n:        m = min(chunk, n - done)        gap = sampler(rng, m)            # top-1 minus top-2 logit gap        err = rng.normal(0.0, sigma, m)  # numerical error on that gap        flips += np.count_nonzero(gap + err < 0)        done += m    return flips / n

    For an exponential and a half-normal gap, the simulation matches the formula within 4% for σ up to 0.03, and the half-normal stays within 2% even at σ = 0.3 (Figure 2). It breaks when the noise is as large as the typical gap: at σ = 1 the exponential case flips 24% of the time against a predicted 40%, because the density is no longer flat over the region that matters.

    A third case shows the formula’s key assumption, f(0)>0f(0) > 0f(0)>0. If the gap density vanishes at zero, the flip rate grows like σ2σ²σ2 instead (it matches σ2/4σ²/4σ2/4 within 8% for σ up to 0.06), and noise matters far less. So which regime does a real model live in?

    Lines: the formula. Dots: simulation, 10 million samples per point.

    A real model

    I ran GPT-2 on 256 windows of 128 tokens from WikiText-2 and compared a bf16bf16bf16 forward pass with an fp32fp32fp32 one. This is a stand-in for realistic serving noise: my laptop can’t reproduce the batch-size effect of a loaded GPU server, but it can measure the logit error a lower-precision kernel produces. I run the transformer body in bf16bf16bf16and the output projection in fp32fp32fp32, so the final logits aren’t rounded to bf16bf16bf16‘s coarse grid.

    import copy, numpy as np, torchfrom datasets import load_datasetfrom transformers import AutoModelForCausalLM, AutoTokenizerm32 = AutoModelForCausalLM.from_pretrained("gpt2").eval().float()m16 = copy.deepcopy(m32).to(torch.bfloat16)W = m32.get_output_embeddings().weight   # fp32 output matrix, both runstok = AutoTokenizer.from_pretrained("gpt2")data = load_dataset("wikitext", "wikitext-2-raw-v1", split="test")text = "nn".join(t for t in data["text"] if t.strip())ids = torch.tensor(tok(text)["input_ids"])n = len(ids) // 128ids = ids[: n * 128].reshape(n, 128)pick = torch.randperm(n, generator=torch.Generator().manual_seed(0))[:256]ids = ids[pick]

    For every position I keep the fp32fp32fp32 top-20 logits, the gap, the error on that gap in the bf16bf16bf16 run, and whether the argmax flipped. (Plotting code omitted.)

    gaps, deltas, flips, top20 = [], [], [], []with torch.no_grad():    for x in ids.split(8):                 # x: [8, 128] token ids        h32 = m32.base_model(x).last_hidden_state[:, 4:]        h16 = m16.base_model(x).last_hidden_state[:, 4:].float()        ref = (h32 @ W.T).reshape(-1, W.shape[0])        alt = (h16 @ W.T).reshape(-1, W.shape[0])        v, i = ref.topk(20, dim=1)         # fp32 top-20 logits        a = alt.gather(1, i)               # same tokens, bf16 run        gap = v[:, 0] - v[:, 1]        gaps.append(gap.numpy())        deltas.append(((a[:, 0] - a[:, 1]) - gap).numpy())        flips.append((alt.argmax(1) != i[:, 0]).numpy())        top20.append(v.numpy())gap, delta, flip, top20 = map(np.concatenate,                              (gaps, deltas, flips, top20))
    # density of the gap at 0 (per logit), and the error near tiesf0 = np.mean(gap < 0.25) / 0.25err = np.abs(delta[gap < 0.5]).mean()print(f"f(0) = {f0:.3f}, error near ties = {err:.4f}")print(f"flip rate: measured {flip.mean():.4f}, "      f"predicted {f0 * err / 2:.4f}")# inject noise of known size into the real logitsfor s in np.logspace(-3, 0.5, 8):    noise = np.random.randn(*top20.shape) * s    print(f"{s:.4f}", np.mean((top20 + noise).argmax(1) != 0))

    The density of the gap at zero is f(0)≈0.83f(0) ≈ 0.83f(0)≈0.83 per logit, and 9%9%9% of positions have a gap below 0.1 logits. The bf16bf16bf16 error on the gap near ties averages 0.020.020.02 logits, so the formula predicts a flip rate of 0.010.01 0.01. I measured 0.010.010.01. Figure 3 (right) repeats the comparison across noise levels by injecting Gaussian noise of known size into the real logits.

    Left: the distribution of the top-two logit gap in GPT-2 near zero. Right: flip rate against noise level. Dots are noise injected into the real logits, the star is bf16 against fp32.

    From tokens to completions

    For one completion, survival is a product of per-token survival probabilities, S(t)≈exp(−Σps)S(t) ≈ exp(−Σ pₛ)S(t)≈exp(−Σps​). If every token carried the same risk ppp, that would be exp(−p⋅t)exp(−p·t)exp(−p⋅t), and half of all completions would have split by token ln 2/p2 / p2/p. Real text isn’t like that.

    Take the Feynman numbers (this is my arithmetic on their published counts). Over the first 102 tokens, 1,000 runs produced zero disagreements. That is about 102,000 token decisions with no flip, so the average per-token flip rate over that stretch is at most about 3×10⁻⁵ (the “rule of three” 95% bound). At token 103 the rate was 8 in 1,000, or 0.8%. That is more than 250 times higher, at a single token. Risk isn’t spread evenly. It sits on a few knife-edge tokens, with long stretches of near-zero hazard in between.

    Prompts also differ in how many knife-edge tokens they contain. To see what that does to survival, I simulated a mean per-token flip rate of 0.005 with fragility varying across prompts (lognormal, log-sd 1):

    p_bar, sigma_ln = 0.005, 1.0t = np.arange(0, 1001)# per-prompt flip probability: lognormal with mean p_barp_i = p_bar * np.exp(rng.normal(0, sigma_ln, 400_000) - 0.5 * sigma_ln**2)uniform = (1 - p_bar) ** t                  # same p for every promptmixture = np.array([np.mean((1 - p_i) ** k) for k in t])  # p varies
    Simulation, not measured data. Left: fraction of completions still identical after t tokens. Right: the same curves on a log scale. A mixture of exponentials has a heavier tail than any single exponential.

    If every prompt had p = 0.005, half of the completions would split by token 139 and 8% would still match at token 500. With the same mean but varying fragility, the median moves to token 210 and 28% still match at token 500. The average flip rate alone doesn’t tell you how long a completion survives. The spread of fragility matters too.

    Where this breaks

    • My noise isn’t your server’s noise. bf16 against fp32 on a laptop reproduces the size of a precision error, not the batch-size mechanism of a loaded GPU cluster. These numbers describe the model, not any hosted API.

    • Two tokens only. The formula ignores the third-best token and assumes the error is symmetric and independent of the gap.

    • A flip isn’t an error. Many flips swap near-equivalent phrasings. The formula says how often outputs differ, not how often they get worse.

    What to do with it

    Predict your flake rate. If your stack has a per-token flip rate p, a completion of L tokens differs between runs with probability about 1−exp(−p⋅L)1 − exp(−p·L)1−exp(−p⋅L). With p = 10⁻³ and 500-token answers, that is about 39%. Measure f(0) on your model and the noise on your stack, and you can estimate it before shipping an eval.

    Don’t assert exact equality on temperature-0 outputs. Compare with a tolerance, or grade semantically. If you control inference, batch-invariant kernels remove the effect for a throughput price. If you don’t, assume two identical calls can differ.

    Takeaways

    1. Temperature 0 is deterministic in the mathematics, not in the arithmetic.

    2. Only near-ties matter: the per-token flip probability is about f(0)⋅E∣Δ∣/2f(0) · E|Δ| / 2f(0)⋅E∣Δ∣/2 the crowding of ties times the noise level.

    3. Risk is spiky. A few knife-edge tokens carry most of it, which is why completions stay identical for a stretch and then split.

    ···

    Sources and credits

    1. He, Horace and Thinking Machines Lab, “Defeating Nondeterminism in LLM Inference”, Thinking Machines Lab: Connectionism, September 10, 2025. https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/ (source of the Feynman experiment and its counts, the batch-invariance explanation, and the serving timings).

    2. Messina, Alberto and Scotta, Stefano, “Introducing Background Temperature to Characterise Hidden Randomness in Large Language Models”, Transactions on Machine Learning Research, 2026. https://openreview.net/forum?id=bz0he4bARF

    3. Radford et al., “Language Models are Unsupervised Multitask Learners” (GPT-2), 2019, and Merity et al., “Pointer Sentinel Mixture Models” (WikiText-2), 2016, both used through Hugging Face Transformers and Datasets.

    4. The flip-probability formula, simulations, real-model experiment, figures and the arithmetic on the Feynman counts are the author’s own.

    Deterministic isnt Temperature
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Previous ArticleAI disqualification yields new Nikon Small World in Motion winner
    Next Article Danu Robotics’ fight to build a better recycling robot
    • Website

    Related Posts

    AI Tools

    Can TypeSafe’s Jev Make AI Agents Safer Without Another LLM?

    AI Tools

    Where Does the Money Go Across Long-Running Coding Agents?

    AI Tools

    Build Your First AI Agent with One Tool Call

    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    One with the world? A new look at brains transformed by psychedelics.

    0 Views

    AI Is Getting Really Good at Messing With Cybercriminals

    0 Views

    Danu Robotics’ fight to build a better recycling robot

    0 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Latest Reviews
    AI Tutorials

    Quantization from the ground up

    AI Tools

    David Sacks is done as AI czar — here’s what he’s doing instead

    AI Reviews

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Most Popular

    One with the world? A new look at brains transformed by psychedelics.

    0 Views

    AI Is Getting Really Good at Messing With Cybercriminals

    0 Views

    Danu Robotics’ fight to build a better recycling robot

    0 Views
    Our Picks

    Quantization from the ground up

    David Sacks is done as AI czar — here’s what he’s doing instead

    Judge sides with Anthropic to temporarily block the Pentagon’s ban

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Contact Us
    • Terms & Conditions
    • Privacy Policy
    • Disclaimer

    © 2026 ainewstoday.co. All rights reserved. Designed by DD.

    Type above and press Enter to search. Press Esc to cancel.