What you will be able to explain
Explain how a selected token changes the context for the next step, predict how temperature changes a distribution, and name one reason a fluent answer can still be wrong.
The essential path is the explanation, the five-step activity and the quiz. The presentation and technical notes are optional. You can mark the lesson complete at any time.
Start with the mechanism
A token is a piece of text defined by a tokenizer, not necessarily a word. A distribution assigns probabilities across the possible next tokens. The context is the input available for this step.
Three steps to keep in mind
- 1
Read the context
Instructions, the request, and previously generated tokens form the current input.
- 2
Score possible continuations
The model assigns scores to tokens; decoding turns those scores into a choice.
- 3
Add and repeat
The chosen token joins the context. The next prediction uses this extended input, until a stop rule or output limit is reached.
Try it yourself
Follow these five steps in order. Everything in the lab runs locally; no real AI model is called.
1 · Predict
Keep the unfinished sentence ‘The cat drinks’ in mind. If we raise temperature from 1 to 2, will the leading candidate gain or lose probability? Say your prediction in one sentence before touching the slider.
2 · Manipulate
- Use ‘The cat drinks…’, seed 42 and temperature 1. If you have already generated text, press Reset: it clears the sequence but keeps your settings. Leave autoplay off.
- Before generating anything, read the percentages for milk, broth and juice. Move only Temperature from 1 to 2. Keep the scenario, seed and unfinished context unchanged; compare the same three candidates.
- Now generate one token. Notice that it joins the context and that the next set of candidates changes. Generate once more to reach the end of this short prepared branch.
- Press Generate A and B. Both runs start from the same scenario with the same seed: A uses temperature 0.4 and B uses your current setting. This comparison does not continue from the on-screen partial sentence.
Local lab
Token-by-token generation
Inspect a prepared distribution, then sample a reproducible sequence. Temperature is only one part of a real generation pipeline.
Controls
Current distribution
Current contextThe cat drinks
- milk68 %
- broth20 %
- juice12 %
Normalized total: 100 %. Individual percentages may differ slightly because of rounding.
Ready. Choose a scenario, then generate one token or start autoplay.
Compare two generations
Same scenario and seed: A uses a low temperature of 0.4; B uses your current setting (1).
Generation A · T 0.4
Not generated yet.
Generation B · T 1
Not generated yet.
What this simulation does not show
- The starting candidates and probabilities are prepared; the slider recalculates their weights.
- There is no neural network, attention mechanism, or real tokenizer.
- A real system may apply top-k, top-p, penalties, and additional stop rules.
- A token probability is not the probability that a claim is true.
- The seed is reproducible only inside this lab.
3 · Observe
At the initial context, milk starts at 68%. At temperature 2 its share falls while the other two shares rise. Compare the displayed percentages, not just the sampled word. A and B may still produce the same sentence: two samples are not a measurement of long-run frequencies.
4 · Explain
Without copying the summary, explain to someone else: what changed when you moved the slider, what changed after selecting a token, and why could A and B still match? Say it aloud or jot it down on paper; nothing is collected here.
After your own explanationCompare with one possible explanation
Raising temperature spread probability more evenly across the existing candidates; it did not add new knowledge or words to the vocabulary. Selecting a token extended the context, so the next prediction used a different input. Different distributions can still select the same token, so different temperatures need not produce different sentences.
5 · Qualify
Open “Token-by-token generation”Check the model in your head
Six short questions, each with an explanation. You can retry or skip the quiz; your best score stays on this device and is shared between languages.
No tricks, just explanations
Check your understanding
Nothing is locked by this quiz. Use mistakes to refine your explanation.
Keep these three ideas
Go deeper when you need it
The essential path is complete. Open only the resources you need; none are required to finish.
Technical detailWhat the short version leaves out
A tokenizer encodes text as numerical IDs before the model processes it. Tokens can be words, word fragments, punctuation or other pieces; different tokenizers can split the same text differently.
A model produces scores called logits over its vocabulary. Softmax converts scores to probabilities. A positive temperature T typically rescales logits before normalization: pᵢ(T) = exp(zᵢ / T) / Σⱼ exp(zⱼ / T).
Temperature changes how concentrated the distribution is. Lower values usually favor leading candidates; higher values give secondary candidates more weight.
Sampling uses these weights for a random draw. Greedy decoding chooses a highest-scoring candidate instead. A changing roulette wheel can help picture sampling, but no physical roulette exists inside a model.
Real systems may also use top-k, top-p, penalties, stop rules, tools, and cached intermediate calculations.
Pretraining often learns to predict tokens in text; later training can improve instruction-following and other behaviors. None of this makes next-token probability an automatic truth test. These explanations concern the common autoregressive case, not every possible model architecture.
Reusable teaching resourceOpen the eight-slide presentation
Use this as a recap or a teaching outline. Download the complete Markdown file for reuse without a network request.
Sources and further reading
References checked for this English lesson on . Publication years are listed separately. These explain the mechanisms, not a current ranking of products; real systems and documentation evolve.
- Attention Is All You Need — Ashish Vaswani et al. (2017) (opens a new tab)Introduces the Transformer. Sections 3 and 3.4 describe autoregressive decoding and next-token probabilities. Its original encoder–decoder architecture is not a description of every modern LLM.
- Language Models are Few-Shot Learners — Tom B. Brown et al. (2020) (opens a new tab)Describes GPT-3 as an autoregressive language model and reports capabilities and limitations. A concrete model example, not a promise about every present-day system.
- SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing — Taku Kudo and John Richardson (2018) (opens a new tab)Primary description of subword tokenization, published at EMNLP 2018. Supports the distinction between text, words and tokenizer-defined pieces; our lab runs no real tokenizer.
- The Curious Case of Neural Text Degeneration — Ari Holtzman et al. (2020) (opens a new tab)Studies how decoding and sampling affect generated text. ICLR 2020 publication; the preprint first appeared in 2019. Quality and diversity are distinct from factual truth.
- TruthfulQA: Measuring How Models Mimic Human Falsehoods — Stephanie Lin, Jacob Hilton and Owain Evans (2022) (opens a new tab)ACL 2022 benchmark showing that tested models can imitate human falsehoods. Evidence that fluency is insufficient, not a current accuracy ranking or a claim that models are always wrong.
Your local progress
Status: Not started
Progress stays on this device.