What you will be able to explain
Distinguish text, tokens and token IDs; show how two vocabularies reconstruct the same text with different sequence lengths, and explain one practical consequence without ranking model quality by token count.
The essential path is the explanation, the five-step activity and the quiz. The presentation and technical notes are optional. You can mark the lesson complete at any time.
Start with the mechanism
A token is a piece of text defined by a tokenizer, not necessarily a word. A distribution assigns probabilities across the possible next tokens. The context is the input available for this step.
Three steps to keep in mind
- 1
Normalize the text
Some tokenizers standardize spaces, case, or Unicode forms before splitting.
- 2
Find vocabulary pieces
The same word may become one token, several subwords, or byte-level pieces.
- 3
Use numerical IDs
Each token maps to an identifier understood only with that tokenizer and version.
Try it yourself
Follow these five steps in order. Everything in the lab runs locally; no real AI model is called.
1 · Predict
For ‘Models learn.’, will the compact and fragmented profiles need the same number of tokens? Predict what can change while the text stays exactly the same.
2 · Manipulate
- Select ‘An English sentence’. Keep this text fixed and leave identifiers hidden. Compare Compact profile with Fragmented profile: only the prepared vocabulary changes between the two columns.
- Read each column from left to right, joining the pieces. The ␠ symbol marks a space; it is a display aid, not a literal character added to the text. Count the pieces in each column.
- Turn on ‘Show prepared identifiers’. Compare the pieces and their IDs again. This changes only the display, not the segmentation or the original text.
- For a second controlled comparison, select ‘A long word’ and compare the two profiles again while keeping that word fixed. You can also inspect ‘Text and emoji’ or ‘A line of code’ one at a time.
Compare two segmentations
One text, several tokenizations
Change the example and see how two prepared vocabularies represent exactly the same text with different boundaries.
Models learn.Fictional vocabulary
Compact profile
This prepared vocabulary contains several frequent units from the sentence.
- Models
- ␠learn
- .
Fictional vocabulary
Fragmented profile
This prepared vocabulary needs more fragments to produce exactly the same text.
- Mod
- els
- ␠le
- arn
- .
A vocabulary may contain a frequent whole word or only fragments that can reconstruct it.
What this demonstration does not show
- Profiles and identifiers are fictional and prepared locally.
- No commercial or open-source tokenizer runs here.
- Real boundaries depend on the algorithm, vocabulary, version, normalization, and exact text.
- Fewer tokens do not automatically imply a better model or answer.
3 · Observe
Both sentence profiles reconstruct ‘Models learn.’, but the compact profile uses 3 tokens and the fragmented profile uses 5. The long word uses 3 versus 5 pieces too. Showing IDs does not change these counts. No profile has added or removed information from the original text.
4 · Explain
Without looking back, explain why the same sentence can occupy different amounts of a token budget, and why an ID is not a meaning or a probability. Say it aloud or write it on paper; nothing is collected.
After your own explanationCompare with one possible explanation
Each vocabulary offers different reusable pieces. Joining either sequence recovers the same text, but more pieces occupy more token positions. An ID indexes a piece within a particular vocabulary; it is not a truth score, a character count or a universal concept number.
5 · Qualify
Open “One text, several tokenizations”Check the model in your head
Six short questions, each with an explanation. You can retry or skip the quiz; your best score stays on this device and is shared between languages.
No tricks, just explanations
Check your understanding
Nothing is locked by this quiz. Use mistakes to refine your explanation.
Keep these three ideas
Go deeper when you need it
The essential path is complete. Open only the resources you need; none are required to finish.
Technical detailWhat the short version leaves out
Subword tokenization balances a finite vocabulary against reasonably short sequences.
Unicode, whitespace, rare words, and languages can all change how much text fits into the same token budget.
Text, token pieces and IDs are different things. An ID indexes a vocabulary entry; the model uses it to look up a learned numerical representation. It is neither a universal meaning nor a probability. Decoding maps generated IDs back to displayable text, subject to the tokenizer’s rules and normalization.
BPE, WordPiece and unigram methods differ. A box of reusable pieces is an analogy, not an instruction to always choose the longest piece. Some systems also add special tokens for roles or boundaries, while byte-based models need no conventional subword tokenizer.
More tokens can affect processing, latency and token-based pricing, depending on the architecture and service. Fewer tokens alone do not establish better coverage, robustness or model quality.
Reusable teaching resourceOpen the eight-slide presentation
Use this as a recap or a teaching outline. Download the complete Markdown file for reuse without a network request.
Sources and further reading
References checked for this English lesson on . Publication years are listed separately. These explain the mechanisms, not a current ranking of products; real systems and documentation evolve.
- Neural Machine Translation of Rare Words with Subword Units — Rico Sennrich, Barry Haddow and Alexandra Birch (2016) (opens a new tab)Introduces BPE-based subword units for rare-word translation. Supports reusable fragments and vocabulary trade-offs, not the exact fictional boundaries displayed here.
- SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing — Taku Kudo and John Richardson (2018) (opens a new tab)Describes subword tokenization trained directly from raw text, with BPE and unigram options. Supports the distinction between text and vocabulary-defined units.
- ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte Models — Linting Xue et al. (2022) (opens a new tab)Studies byte-based text models. ‘Token-free’ here contrasts with conventional subword tokenizers; the model still processes numerical units, not unencoded human language.
- Tokenizer Choice For LLM Training: Negligible or Crucial? — Mehdi Ali et al. (2024) (opens a new tab)Controlled study of tokenizer choices for monolingual and multilingual models. Supports considering downstream performance and coverage rather than judging quality by token count alone.
Your local progress
Status: Not started
Progress stays on this device.