Understand LLMs · Lesson 03

Why can two answers be different?

How can we separate randomness from changes in context, model, or environment?

Level
Beginner
Core time (estimate)
12 min
Updated
11 September 2026
Prerequisites
Progress
Not started

What you will be able to explain

Diagnose a changed input separately from sampling and numerical variation, explain why a seed is conditional rather than a universal guarantee, and propose a one-factor comparison.

The essential path is the explanation, the five-step activity and the quiz. The presentation and technical notes are optional. You can mark the lesson complete at any time.

Start with the mechanism

A token is a piece of text defined by a tokenizer, not necessarily a word. A distribution assigns probabilities across the possible next tokens. The context is the input available for this step.

Three steps to keep in mind

  1. 1

    Compare the inputs

    Hidden instructions, history, retrieved documents, and formatting all belong to the context.

  2. 2

    Compare decoding

    Temperature, top-k, top-p, penalties, and the seed can change token selection.

  3. 3

    Compare the system

    Model versions, libraries, hardware, and parallel execution may affect reproducibility.

Try it yourself

Follow these five steps in order. Everything in the lab runs locally; no real AI model is called.

1 · Predict

If the prompt, model and seed stay the same, can a variable execution environment still weaken repeatability? Predict the diagnosis before changing a control.

2 · Manipulate

  1. Start with Prompt and context, Model and version, and External data or tools set to Same; use Greedy decoding, the same seed and a Controlled and identical environment. Reset conditions restores this baseline.
  2. Change only Execution environment to Variable or unknown. Compare the diagnosis with the controlled baseline. Keep Greedy decoding and all inputs unchanged.
  3. Return the environment to Controlled and identical. In a separate comparison change Decoding strategy to Sampling, keeping the seed Same. Then change only the seed to Different and compare again.
  4. Reset conditions. Change only External data or tools to Changed. Read why this is a different category from sampling under otherwise identical conditions.

Isolate what changed

Reproducibility diagnosis

Compare two fictional runs and change one factor at a time to identify possible causes of a different answer.

Prepared reasoning

Teaching diagnosis

The conditions favor repeatability

In this controlled setup, greedy decoding should select the highest-scoring token at every step.

  • the prompt, model, and external data are identical
  • selection is greedy
  • the execution environment is declared controlled

Keep in mind: This diagnosis states an expectation for a controlled system, never a universal guarantee across services, versions, or hardware.

Limits of this diagnosis
  • No model or remote service is run.
  • The result classifies prepared conditions; it does not predict a measured probability of divergence.
  • A real service may hide settings, change versions, or add steps unknown to this grid.
  • Different wording may express the same information, and identical wording can still be wrong.

3 · Observe

The controlled baseline says ‘The conditions favor repeatability’. Changing only the environment gives ‘Variation remains possible’, even with Greedy decoding. Controlled sampling with the same seed favors repeatability; changing just that seed gives ‘Variation is likely’. Changing tool data instead reports that the runs do not start from the same conditions.

4 · Explain

Explain, in your own words, which comparison isolated numerical uncertainty and which changed the actual input. Why does a fixed seed not solve both? Say it aloud or write a short note; nothing is collected.

After your own explanationCompare with one possible explanation

Greedy selection has no random draw, but numerical changes can alter close scores and therefore the selected token. A seed controls a pseudo-random sequence only within a compatible pipeline. Tool data changes the context itself, so the comparison no longer isolates randomness. After one different token, later contexts and predictions can diverge too.

5 · Qualify

Open “Reproducibility diagnosis”

Check the model in your head

Six short questions, each with an explanation. You can retry or skip the quiz; your best score stays on this device and is shared between languages.

No tricks, just explanations

Check your understanding

Nothing is locked by this quiz. Use mistakes to refine your explanation.

Question 1 of 6
Runs with the same prompt can produce different answers when sampling is used.

Keep these three ideas

Go deeper when you need it

The essential path is complete. Open only the resources you need; none are required to finish.

Technical detailWhat the short version leaves out

Greedy decoding removes the random draw at each local step, but it does not guarantee identical results across every infrastructure.

To diagnose variation, change one factor at a time and record the complete configuration.

Top-k restricts sampling to k leading candidates; top-p selects a dynamic set reaching a cumulative probability threshold. Temperature rescales their relative weights. These controls do different jobs.

One different token extends the two contexts differently, so the following distributions can diverge. Branching routes are an analogy, not a literal map inside the model.

Record full instructions, history, model and tokenizer versions, tool results, decoding settings, date, libraries and hardware where available. Repeat runs before estimating variation, and compare meaning and evidence as well as exact text.

Self-consistency deliberately samples several reasoning paths and aggregates their final answers in studied tasks. Diversity can be useful, but repeated agreement is not independent evidence that an answer is true.

Reusable teaching resourceOpen the eight-slide presentation

Use this as a recap or a teaching outline. Download the complete Markdown file for reuse without a network request.

When the presentation has focus, use the left and right arrows to change slides.

The question

Why can two answers differ?

The visible prompt is not the whole system. Context, decoding, model version, tool data and numerical execution can all matter.

Slide 1: Why can two answers differ?

1 / 8

Sources and further reading

References checked for this English lesson on . Publication years are listed separately. These explain the mechanisms, not a current ranking of products; real systems and documentation evolve.

Primary sources

Your local progress

Status: Not started

Progress stays on this device.