Understand LLMs · Lesson 02

What is a token?

How does text become a sequence of numerical units a model can process?

Level
Beginner
Core time (estimate)
11 min
Updated
11 September 2026
Prerequisites
Progress
Not started

What you will be able to explain

Distinguish text, tokens and token IDs; show how two vocabularies reconstruct the same text with different sequence lengths, and explain one practical consequence without ranking model quality by token count.

The essential path is the explanation, the five-step activity and the quiz. The presentation and technical notes are optional. You can mark the lesson complete at any time.

Start with the mechanism

A token is a piece of text defined by a tokenizer, not necessarily a word. A distribution assigns probabilities across the possible next tokens. The context is the input available for this step.

Three steps to keep in mind

  1. 1

    Normalize the text

    Some tokenizers standardize spaces, case, or Unicode forms before splitting.

  2. 2

    Find vocabulary pieces

    The same word may become one token, several subwords, or byte-level pieces.

  3. 3

    Use numerical IDs

    Each token maps to an identifier understood only with that tokenizer and version.

Try it yourself

Follow these five steps in order. Everything in the lab runs locally; no real AI model is called.

1 · Predict

For ‘Models learn.’, will the compact and fragmented profiles need the same number of tokens? Predict what can change while the text stays exactly the same.

2 · Manipulate

  1. Select ‘An English sentence’. Keep this text fixed and leave identifiers hidden. Compare Compact profile with Fragmented profile: only the prepared vocabulary changes between the two columns.
  2. Read each column from left to right, joining the pieces. The ␠ symbol marks a space; it is a display aid, not a literal character added to the text. Count the pieces in each column.
  3. Turn on ‘Show prepared identifiers’. Compare the pieces and their IDs again. This changes only the display, not the segmentation or the original text.
  4. For a second controlled comparison, select ‘A long word’ and compare the two profiles again while keeping that word fixed. You can also inspect ‘Text and emoji’ or ‘A line of code’ one at a time.

Compare two segmentations

One text, several tokenizations

Change the example and see how two prepared vocabularies represent exactly the same text with different boundaries.

Prepared demonstration
Original textModels learn.

Fictional vocabulary

Compact profile

3 tokens

This prepared vocabulary contains several frequent units from the sentence.

  1. Models
  2. ␠learn
  3. .

Fictional vocabulary

Fragmented profile

5 tokens

This prepared vocabulary needs more fragments to produce exactly the same text.

  1. Mod
  2. els
  3. ␠le
  4. arn
  5. .

A vocabulary may contain a frequent whole word or only fragments that can reconstruct it.

What this demonstration does not show
  • Profiles and identifiers are fictional and prepared locally.
  • No commercial or open-source tokenizer runs here.
  • Real boundaries depend on the algorithm, vocabulary, version, normalization, and exact text.
  • Fewer tokens do not automatically imply a better model or answer.

3 · Observe

Both sentence profiles reconstruct ‘Models learn.’, but the compact profile uses 3 tokens and the fragmented profile uses 5. The long word uses 3 versus 5 pieces too. Showing IDs does not change these counts. No profile has added or removed information from the original text.

4 · Explain

Without looking back, explain why the same sentence can occupy different amounts of a token budget, and why an ID is not a meaning or a probability. Say it aloud or write it on paper; nothing is collected.

After your own explanationCompare with one possible explanation

Each vocabulary offers different reusable pieces. Joining either sequence recovers the same text, but more pieces occupy more token positions. An ID indexes a piece within a particular vocabulary; it is not a truth score, a character count or a universal concept number.

5 · Qualify

Open “One text, several tokenizations”

Check the model in your head

Six short questions, each with an explanation. You can retry or skip the quiz; your best score stays on this device and is shared between languages.

No tricks, just explanations

Check your understanding

Nothing is locked by this quiz. Use mistakes to refine your explanation.

Question 1 of 6
A token always corresponds to one complete word.

Keep these three ideas

Go deeper when you need it

The essential path is complete. Open only the resources you need; none are required to finish.

Technical detailWhat the short version leaves out

Subword tokenization balances a finite vocabulary against reasonably short sequences.

Unicode, whitespace, rare words, and languages can all change how much text fits into the same token budget.

Text, token pieces and IDs are different things. An ID indexes a vocabulary entry; the model uses it to look up a learned numerical representation. It is neither a universal meaning nor a probability. Decoding maps generated IDs back to displayable text, subject to the tokenizer’s rules and normalization.

BPE, WordPiece and unigram methods differ. A box of reusable pieces is an analogy, not an instruction to always choose the longest piece. Some systems also add special tokens for roles or boundaries, while byte-based models need no conventional subword tokenizer.

More tokens can affect processing, latency and token-based pricing, depending on the architecture and service. Fewer tokens alone do not establish better coverage, robustness or model quality.

Reusable teaching resourceOpen the eight-slide presentation

Use this as a recap or a teaching outline. Download the complete Markdown file for reuse without a network request.

When the presentation has focus, use the left and right arrows to change slides.

The question

What is a token?

A tokenizer defines units for encoding text. A token is not necessarily a complete word or one visible character.

Slide 1: What is a token?

1 / 8

Sources and further reading

References checked for this English lesson on . Publication years are listed separately. These explain the mechanisms, not a current ranking of products; real systems and documentation evolve.

Primary sources

Your local progress

Status: Not started

Progress stays on this device.