Glossary · 5 minute read
What Is Top-p Sampling? Nucleus Sampling Explained
Top-p sampling, or nucleus sampling, restricts each token choice to the smallest set of candidates whose probabilities sum to p, then samples from that set. It adapts to the model's confidence: few candidates where the model is certain, many where it is not. It is usually preferred over top-k for that reason.
Sampling parameters are usually left at their defaults, which were chosen for open-ended chat rather than for extraction, classification, or agent tool selection. The result is production systems with avoidable variability in exactly the places variability is unwanted. This explainer covers what top-p does, how it relates to temperature and top-k, and how to set it deliberately. It complements deterministic ai outcomes and why determinism matters in ai systems, and reflects FISTA Solutions' approach in AI agents delivery.
What does top-p actually do?
At each generation step the model produces a probability distribution over its entire vocabulary. Top-p sorts candidates by probability, accumulates them until the running total reaches p, discards the rest, and samples from what remains.
At p of 0.9, if the model is highly confident the set may contain two or three tokens; if it is uncertain the set may contain hundreds. The candidate pool adapts to confidence, which is the property that makes nucleus sampling useful.
| Parameter | What it controls | Adapts to confidence |
|---|---|---|
| Top-p | Cumulative probability cutoff | Yes |
| Top-k | Fixed candidate count | No |
| Temperature | Distribution sharpness | N/A |
| Repetition penalties | Token reuse | No |
| Seed | Sampling randomness source | N/A |
Why is top-k weaker?
Because a fixed count ignores the distribution. Where the model is confident, top-k of 50 admits 49 low-probability candidates that should not be considered. Where the model is genuinely uncertain across hundreds of plausible continuations, the same setting truncates arbitrarily.
Top-p handles both cases with one parameter. Top-k remains available and is occasionally useful as a hard bound alongside top-p, but it is rarely the right primary control.
How does temperature interact?
Temperature is applied first, reshaping the distribution. Below one it sharpens â the most probable tokens become more probable â and above one it flattens, raising the chance of unlikely tokens. Top-p then truncates whatever distribution results.
Because both affect randomness through different mechanisms, adjusting them simultaneously produces effects that are hard to predict. The common practice is to vary temperature and leave top-p near one, or vary top-p and keep temperature fixed.
What settings suit which tasks?
Extraction, classification, structured output, and tool selection want the most probable continuation. Temperature at or near zero with top-p near one is appropriate: randomness there produces inconsistency with no compensating benefit.
Drafting, brainstorming, and content variation can tolerate and sometimes benefit from higher randomness. Even then, production systems usually want less than chat defaults provide, because reviewers value consistency.
Does zero temperature make output reproducible?
No. Greedy decoding removes sampling randomness, which is one source of variation, but not the only one. Floating-point non-determinism, batching effects on shared inference infrastructure, and provider-side model updates can all produce different outputs for identical inputs.
Systems that genuinely require reproducibility need caching, a recorded output, or a deterministic path â not a temperature setting. See deterministic ai outcomes.
Why does this matter for evaluation?
Because sampling settings change the output distribution, so evaluation results are only valid for the configuration they were measured under. An accuracy figure obtained at temperature zero does not describe behaviour at temperature 0.7.
Sampling parameters therefore belong in the evaluation configuration alongside the model version and the prompt version, and a change to any of them invalidates prior results.
What about agent tool selection?
It deserves particular attention. An agent choosing which tool to call is making a discrete decision where the second-most-probable option is frequently wrong in a consequential way. Randomness in that choice produces occasional inexplicable behaviour that is extremely hard to debug, because it does not reproduce.
Tool selection should run at minimal randomness regardless of what the generation steps use.
How should teams configure this?
Deliberately, per task type, recorded in configuration rather than scattered through code. Most systems need two or three named profiles â deterministic, balanced, creative â applied by task rather than a single global setting inherited from a framework default.
What about structured output modes?
Where a provider offers constrained decoding against a schema, it changes the calculation. Constraining generation to valid JSON or to an enumerated set removes whole categories of error that sampling settings only reduce, and it usually makes low-temperature configuration less critical for format correctness â though not for content consistency.
Where constrained decoding is unavailable, low randomness plus schema validation in application code is the practical substitute, with a retry on validation failure rather than a hope that the next call parses.
How should changes be rolled out?
Like model changes, because they are behaviour changes. A sampling adjustment that improves one task's consistency can degrade another's usefulness, and the effect is invisible without evaluation. Treating sampling configuration as versioned, evaluated, and deployed rather than as a knob anyone can turn in production is the difference between a tuned system and an unexplained regression.
How FISTA Solutions helps
FISTA Solutions configures sampling per task type rather than globally, uses minimal randomness for extraction and tool selection, records sampling parameters as part of evaluation configuration, and builds real determinism through caching where reproducibility is required, through AI agents, AI enablement, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To remove avoidable variability from your AI systems, message FISTA on WhatsApp, or read why determinism matters in ai systems.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01How does top-p differ from top-k?
Top-k always considers a fixed number of candidates regardless of the distribution. Top-p considers however many are needed to reach the probability threshold, which is few when the model is confident and many when it is uncertain. That adaptivity is why it is generally preferred.
02How does it interact with temperature?
Temperature reshapes the probability distribution before truncation, flattening it to increase randomness or sharpening it to reduce it. Top-p then truncates the tail. Adjusting both at once makes the effect hard to reason about, so most teams vary one and fix the other.
03What settings suit extraction tasks?
Low randomness: temperature at or near zero, with top-p at or close to one so it has little effect. Extraction, classification and structured output want the most probable continuation, not variety, and randomness there produces inconsistency without any benefit.
04Does setting temperature to zero guarantee reproducibility?
No. It makes sampling greedy, which removes one source of variation, but batching on shared inference infrastructure, floating-point differences, and provider-side model updates can still produce different outputs for identical inputs. Real reproducibility needs caching or a recorded output, not a temperature setting.
05Why do sampling settings matter for evaluation?
Because they change the output distribution. An evaluation run at one temperature does not predict behaviour at another, so sampling parameters must be fixed and recorded as part of the evaluation configuration alongside the model version.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.