← all articles

Temperature, top-p and sampling explained

If you have ever called a language model through an API, you have probably seen a few odd parameters sitting next to the prompt: temperature, top_p, sometimes top_k. Most people leave them at the defaults. That is fine until the model gives a different answer to the same question, or writes something bland, or makes something up with total confidence. Then these dials matter.

I run a few small operations in Singapore that lean on models for drafting, classification and data cleanup, so I touch these settings often. This article explains what they do in plain language, with no hard maths. By the end you should know which dial to reach for when you want a model steadier or looser, and which popular advice to ignore.

what it is

A language model does not write a sentence in one go. It writes one token at a time. A token is a chunk of text, roughly a short word or part of a word. At every step the model produces a score for every token in its vocabulary, which can run to tens of thousands of options or more. Those scores get turned into probabilities: “the” might get 40%, “a” 25%, “this” 10%, and a long tail of thousands of other tokens share the rest.

Sampling is the step where the system picks one token from that list of probabilities. Temperature, top-p and top-k are the settings that shape how that pick is made:

  • temperature: a number that flattens or sharpens the probability list before the pick.
  • top-p (also called nucleus sampling): cut the list down to the smallest group of top tokens whose probabilities add up to p, then pick only from that group.
  • top-k: keep only the k most likely tokens and pick from those.

None of these change what the model knows. They change how adventurous the pick is, given what the model already thinks is likely. I will come back to that distinction in the misconceptions section.

how it works

Start with the simplest possible approach. Always pick the single most probable token. This is called greedy decoding. It is predictable, but it tends to produce repetitive, flat text, and it can get stuck in loops. The 2019 paper “The Curious Case of Neural Text Degeneration” by Holtzman and colleagues, available on arXiv, showed that always choosing the most likely words produces dull and repetitive output, and that pure randomness produces nonsense. Top-p sampling came out of that work as a middle path.

So instead of always taking the top token, we roll a weighted die. A token with 40% probability gets picked about 40% of the time. That is plain sampling, and it is where temperature comes in.

Temperature. Before the scores become probabilities, the model’s raw scores (called logits) are divided by the temperature. The effect is easy to describe:

  • low temperature (say 0.2): the gap between likely and unlikely tokens grows. The top token gets even more dominant. Output becomes focused and repeatable.
  • temperature 1.0: the probabilities are used as the model produced them. This is the default on most APIs.
  • high temperature (say 1.5): the gap shrinks. Unlikely tokens get a real chance. Output becomes more varied, and also more likely to wander or break.

As temperature approaches zero, you approach greedy decoding. The valid range depends on the provider. OpenAI’s API reference documents temperature between 0 and 2 with a default of 1, while Anthropic’s Messages API documentation documents a range of 0.0 to 1.0 with a default of 1.0. Always check the docs for your specific model, because limits and defaults differ between vendors and versions.

Top-p. Imagine the tokens sorted from most to least likely. Top-p of 0.9 means: walk down the list adding up probabilities until you reach 90%, then throw away everything below that point and pick from what is left (after renormalising). The clever part is that the cut-off adapts. When the model is very sure, say a 95% token, the nucleus may be a single token. When the model is unsure and many tokens are plausible, the nucleus is wide. A fixed top-k cannot do that. It keeps the same number of tokens whether the model is confident or confused.

Top-k. Keep the k highest-probability tokens, discard the rest. Top-k of 40 keeps the best 40 candidates. Some APIs expose it and some do not. The Hugging Face generation strategies guide covers greedy search, top-k, top-p and beam search side by side with code, and is a good place to see them in practice if you self-host.

How they combine. These filters stack. A typical pipeline applies temperature to reshape the probabilities, then trims with top-k or top-p, then samples. Because the settings overlap in what they do, both OpenAI and Anthropic suggest in their docs that you adjust temperature or top-p, but not both at once. Change one, observe, and only then consider the other.

Even at temperature 0, output is not guaranteed to be identical between runs on hosted models. Differences in hardware, batching and floating point arithmetic can shift which token wins a close call. Treat temperature 0 as “as repeatable as the provider can make it”, not as a guarantee.

why it matters

Where I care about these settings in real work:

  • extraction and classification: if I am pulling invoice fields out of messy text or tagging support tickets, I want the same input to give the same label. Low temperature reduces the chance that a borderline ticket gets a different tag on Tuesday than it did on Monday.
  • creative and drafting work: for brainstorming headlines or product names, a bit more temperature gives variety. At low settings I get the same three safe ideas every time; higher settings give me more options to throw away, which is the point.
  • code and structured output: for JSON, SQL or code, one wrong token can break the whole thing. Lower randomness makes malformed output less likely, though it will not fix a model that does not know the answer. If you let a model touch real files, the guardrails in letting a model touch your repository safely matter far more than any sampling setting.
  • testing and debugging: if you cannot tell whether a change to your prompt helped, part of the reason may be that sampling noise is hiding the signal. I cover how to handle that in testing an AI feature before you ship it. Short version: run each test case several times, not once, and fix or lower randomness while comparing prompts.

There is a cost angle too. Sampling settings do not change the price per token, but unpredictable output can trigger retries and manual review, and those cost real money.

common misconceptions

“High temperature makes the model smarter or more creative.” It makes the output more random. Sometimes random looks creative, but the model has not gained any ability. Raise temperature far enough and you get word salad. The knowledge and reasoning are fixed by the model. Temperature only changes how much of the model’s uncertainty shows up in the text.

“Temperature 0 removes hallucinations.” It does not. If the model’s most likely continuation is a wrong fact, greedy decoding will state the wrong fact every time, with full confidence. Lower temperature gives you consistency, and consistency is not accuracy. When answers are wrong because the model is missing the right information, the fix is usually retrieval and better context. I wrote about that in why your RAG answers are wrong.

“Temperature 0 is fully deterministic.” Mostly, but not always, as covered above. Hosted inference involves batching and parallel hardware, and tiny numeric differences can flip a close decision. If you need exact reproducibility, log your outputs instead of trusting regeneration. Some providers offer a seed parameter, which helps but is usually described as best effort.

“You should tune every setting at once.” This is how people end up with a config nobody can explain. Pick one dial, change it in small steps, and test on a fixed set of examples. For most production work I start at the vendor default, then try a lower temperature for anything that needs consistency. I rarely touch top-p or top-k unless I am self-hosting an open model and need tighter control. If you do self-host, the best open-source LLMs you can self-host in 2026 is a reasonable starting point, since local runtimes expose more of these knobs than hosted APIs do.

where to go from here

If this made sense, these are the natural next steps:

  • speed: sampling is part of what makes text generation slow, since tokens come out one at a time. Speculative decoding explained shows how providers get faster output without changing the answer distribution.
  • evaluation: before you tune anything, build a small test set. Testing an AI feature before you ship it walks through how I do that without a big budget.
  • retrieval: if your real problem is wrong answers rather than dull ones, read why your RAG answers are wrong before you touch another sampling setting.
  • the wider index: everything else on the site sits in the blog index, including other beginner explainers.

If you publish content a model helped draft, my sister site The SEO Desk covers how it tends to perform in search, a separate question from how you sampled it.

The main thing I would like you to take away: temperature and top-p are steering dials, not intelligence dials. Use low randomness when you need repeatable output, a little more when you want options, and change one setting at a time. If the answers are wrong rather than merely different, the sampling setting is almost never the real problem.

Written by Xavier Fok

disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-10-01.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →