What you will be able to explain
Explain how input and output compete for a finite context budget, predict which prepared block is removed when capacity shrinks, and distinguish available context from durable memory or reliable use of every passage.
The essential path is the explanation, the five-step activity and the quiz. The presentation and technical notes are optional. You can mark the lesson complete at any time.
Start with the mechanism
A token is a piece of text defined by a tokenizer, not necessarily a word. A distribution assigns probabilities across the possible next tokens. The context is the input available for this step.
Three steps to keep in mind
- 1
Reserve output space
A system must leave enough capacity for the answer it wants to generate.
- 2
Prioritize the input
Instructions and relevant evidence compete with history and supporting material.
- 3
Handle overflow
Applications may truncate, summarize, retrieve selectively, or reject oversized input.
Try it yourself
Follow these five steps in order. Everything in the lab runs locally; no real AI model is called.
1 · Predict
With every block selected, will increasing total capacity from 1400 to 1800 units restore Older conversation? Predict which block changes and whether the answer reserve will change.
2 · Manipulate
- Select 1400 units and check all four blocks. Keep their requested sizes and retention order fixed. Read the input budget, requested units and which block is removed.
- Change only Prepared total capacity to 1800 units. Compare Older conversation, the input budget and the answer reserve with the previous state.
- Return only capacity to 1400 units to test reversibility. For a separate comparison at this capacity, uncheck Retrieved document while leaving the other blocks requested. Observe whether Older conversation now fits.
- Recheck Retrieved document to restore the original request. Explain the trade-off before opening the reference explanation: the lab retains whole blocks in a fixed order, not by understanding their relevance.
Build a limited input
Context-window budget
Enable blocks and change capacity to see what fits after reserving space for the answer.
Prepared plan
1 block removed
Input budget: 1100 units. Requested: 1360 units. Theoretical overflow: 260 units.
Limits of this representation
- The units are not real token counts.
- This strategy keeps blocks in a fixed order; real applications may truncate, summarize, or retrieve differently.
- Fitting in the window does not guarantee that the model will use every piece of information equally well.
3 · Observe
At 1400 total units, 300 are reserved for the answer, leaving 1100 for input. The four requested blocks total 1360; the plan keeps 1000 and removes Older conversation (360). At 1800 total units the input budget becomes 1500 and every requested block fits. The answer reserve stays 300. Removing the retrieved document at 1400 lets the older conversation fit instead.
4 · Explain
Without copying the numbers, explain why adding capacity and removing a document can both restore an older exchange, yet neither guarantees that a real model will use it correctly. Say it aloud or write it down; nothing is collected.
After your own explanationCompare with one possible explanation
The output reserve reduces space available to the input. More capacity increases that input space; removing a large earlier block frees space for a later one under this fixed policy. A block can fit without being relevant or reliably used. Context influences the current inference; it is not automatically permanent knowledge in the model’s parameters.
5 · Qualify
Open “Build a context window”Check the model in your head
Six short questions, each with an explanation. You can retry or skip the quiz; your best score stays on this device and is shared between languages.
No tricks, just explanations
Check your understanding
Nothing is locked by this quiz. Use mistakes to refine your explanation.
Keep these three ideas
Go deeper when you need it
The essential path is complete. Open only the resources you need; none are required to finish.
Technical detailWhat the short version leaves out
Attention lets positions exchange weighted information, but it does not mean every token influences the output equally.
A KV cache speeds up repeated decoding calculations; it is not durable user memory.
Putting a document in context normally changes the current inference, not the model’s learned parameters. An application may keep external memory, but it needs its own retention and retrieval rules and must supply selected material again.
A backpack illustrates a finite budget, not physical compartments inside a model. The lab skips whole blocks in a fixed order; real applications may reject, truncate, summarize or retrieve. A summary or retrieval can lose the detail that made a request meaningful.
Lost-in-the-middle evaluations show position-dependent performance in studied tasks. They do not mean that middle tokens are always deleted. A large window is a capacity, not a promise of perfect recall; relevance, placement and empirical testing remain important.
Reusable teaching resourceOpen the eight-slide presentation
Use this as a recap or a teaching outline. Download the complete Markdown file for reuse without a network request.
Sources and further reading
References checked for this English lesson on . Publication years are listed separately. These explain the mechanisms, not a current ranking of products; real systems and documentation evolve.
- Attention Is All You Need — Ashish Vaswani et al. (2017) (opens a new tab)Introduces the Transformer and attention between sequence positions. Architectural background, not a specification of current chatbot window sizes or persistent memory.
- Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context — Zihang Dai et al. (2019) (opens a new tab)Studies fixed-length context limits and segment-level recurrence. Reusing earlier hidden states in this architecture does not imply that every chatbot remembers previous sessions.
- Lost in the Middle: How Language Models Use Long Contexts — Nelson F. Liu et al. (2024) (opens a new tab)Shows position-dependent performance on multi-document question answering and key-value retrieval. Supports distinguishing window capacity from reliable use of each passage.
- Found in the middle: Calibrating Positional Attention Bias Improves Long Context Utilization — Cheng-Yu Hsieh et al. (2024) (opens a new tab)Investigates positional attention bias and tests a calibration method. Improvements in the studied models and tasks are not a universal cure for context limitations.
Your local progress
Status: Not started
Progress stays on this device.