What you will be able to explain
Diagnose a changed input separately from sampling and numerical variation, explain why a seed is conditional rather than a universal guarantee, and propose a one-factor comparison.
The essential path is the explanation, the five-step activity and the quiz. The presentation and technical notes are optional. You can mark the lesson complete at any time.
Start with the mechanism
A token is a piece of text defined by a tokenizer, not necessarily a word. A distribution assigns probabilities across the possible next tokens. The context is the input available for this step.
Three steps to keep in mind
- 1
Compare the inputs
Hidden instructions, history, retrieved documents, and formatting all belong to the context.
- 2
Compare decoding
Temperature, top-k, top-p, penalties, and the seed can change token selection.
- 3
Compare the system
Model versions, libraries, hardware, and parallel execution may affect reproducibility.
Try it yourself
Follow these five steps in order. Everything in the lab runs locally; no real AI model is called.
1 · Predict
If the prompt, model and seed stay the same, can a variable execution environment still weaken repeatability? Predict the diagnosis before changing a control.
2 · Manipulate
- Start with Prompt and context, Model and version, and External data or tools set to Same; use Greedy decoding, the same seed and a Controlled and identical environment. Reset conditions restores this baseline.
- Change only Execution environment to Variable or unknown. Compare the diagnosis with the controlled baseline. Keep Greedy decoding and all inputs unchanged.
- Return the environment to Controlled and identical. In a separate comparison change Decoding strategy to Sampling, keeping the seed Same. Then change only the seed to Different and compare again.
- Reset conditions. Change only External data or tools to Changed. Read why this is a different category from sampling under otherwise identical conditions.
Isolate what changed
Reproducibility diagnosis
Compare two fictional runs and change one factor at a time to identify possible causes of a different answer.
Teaching diagnosis
The conditions favor repeatability
In this controlled setup, greedy decoding should select the highest-scoring token at every step.
- the prompt, model, and external data are identical
- selection is greedy
- the execution environment is declared controlled
Keep in mind: This diagnosis states an expectation for a controlled system, never a universal guarantee across services, versions, or hardware.
Limits of this diagnosis
- No model or remote service is run.
- The result classifies prepared conditions; it does not predict a measured probability of divergence.
- A real service may hide settings, change versions, or add steps unknown to this grid.
- Different wording may express the same information, and identical wording can still be wrong.
3 · Observe
The controlled baseline says ‘The conditions favor repeatability’. Changing only the environment gives ‘Variation remains possible’, even with Greedy decoding. Controlled sampling with the same seed favors repeatability; changing just that seed gives ‘Variation is likely’. Changing tool data instead reports that the runs do not start from the same conditions.
4 · Explain
Explain, in your own words, which comparison isolated numerical uncertainty and which changed the actual input. Why does a fixed seed not solve both? Say it aloud or write a short note; nothing is collected.
After your own explanationCompare with one possible explanation
Greedy selection has no random draw, but numerical changes can alter close scores and therefore the selected token. A seed controls a pseudo-random sequence only within a compatible pipeline. Tool data changes the context itself, so the comparison no longer isolates randomness. After one different token, later contexts and predictions can diverge too.
5 · Qualify
Open “Reproducibility diagnosis”Check the model in your head
Six short questions, each with an explanation. You can retry or skip the quiz; your best score stays on this device and is shared between languages.
No tricks, just explanations
Check your understanding
Nothing is locked by this quiz. Use mistakes to refine your explanation.
Keep these three ideas
Go deeper when you need it
The essential path is complete. Open only the resources you need; none are required to finish.
Technical detailWhat the short version leaves out
Greedy decoding removes the random draw at each local step, but it does not guarantee identical results across every infrastructure.
To diagnose variation, change one factor at a time and record the complete configuration.
Top-k restricts sampling to k leading candidates; top-p selects a dynamic set reaching a cumulative probability threshold. Temperature rescales their relative weights. These controls do different jobs.
One different token extends the two contexts differently, so the following distributions can diverge. Branching routes are an analogy, not a literal map inside the model.
Record full instructions, history, model and tokenizer versions, tool results, decoding settings, date, libraries and hardware where available. Repeat runs before estimating variation, and compare meaning and evidence as well as exact text.
Self-consistency deliberately samples several reasoning paths and aggregates their final answers in studied tasks. Diversity can be useful, but repeated agreement is not independent evidence that an answer is true.
Reusable teaching resourceOpen the eight-slide presentation
Use this as a recap or a teaching outline. Download the complete Markdown file for reuse without a network request.
Sources and further reading
References checked for this English lesson on . Publication years are listed separately. These explain the mechanisms, not a current ranking of products; real systems and documentation evolve.
- The Curious Case of Neural Text Degeneration — Ari Holtzman et al. (2020) (opens a new tab)ICLR 2020 paper, first posted in 2019. Compares decoding strategies and introduces nucleus sampling; supports distinguishing probability, diversity and output quality.
- Language Models are Few-Shot Learners — Tom B. Brown et al. (2020) (opens a new tab)Demonstrates how instructions and examples in context can affect GPT-3 behavior without gradient updates at inference. A changed context is not the same as retraining the model.
- Self-Consistency Improves Chain of Thought Reasoning in Language Models — Xuezhi Wang et al. (2023) (opens a new tab)ICLR 2023 paper, first posted in 2022. Samples reasoning paths and aggregates final answers on selected benchmarks; agreement is not independent proof of truth.
- Reproducibility — PyTorch contributors (2026) (opens a new tab)Official guidance on seeds and nondeterministic operations. The year refers to the documentation update, not its original publication. The stable documentation can evolve; a seed is not a cross-platform guarantee.
- Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference — Jiayi Yuan et al. (2025) (opens a new tab)Studies precision, hardware and batching as sources of differing LLM outputs, including greedy decoding. Shows a possible mechanism, not that every repeated request must differ.
Your local progress
Status: Not started
Progress stays on this device.