Follow this message

S01 · It starts with a message

Prepared message : Write a message to let someone know I will be late.

Prepared reply — written by our team
Imposed choices, recorded calculations. No live model call.

The same message leaves the application; no answer has been added.

Read the prepared reply (outside the model input)

Hi, I will be a little late. Sorry to keep you waiting!

What am I talking to when I type a message?

You type into an application. It is the interface, not the language model itself.

The reply is authored; model calculations along the imposed path were recorded separately. No live model call.

  1. YouA prepared request.
  2. ApplicationThe screen where you type.
  3. Your messageThe request, not the reply.
The interface is not the model.◇ Teaching diagram · illustrative shapes and distances, not a model measurement.

The reply is authored; model calculations along the imposed path were recorded separately. No live model call.

Show me

Will the reply already be sitting inside this message?

The application prepares a request. Producing the reply is a separate step.

Go deeper

This example follows a remotely hosted assistant. The person, computer and journey are teaching diagrams, not recordings of a real conversation.

Cette scène en français

S02 · The request reaches a server

Prepared message : Write a message to let someone know I will be late.

Does the model receive only the text I can see?

A server is a computer running software. The model is the learned calculation that software uses; it is not the server itself.

The application prepares the input. Open this example's envelope to see exactly what was included.

  1. ApplicationPrepares the input.
  2. ServerThe hosting computer.
  3. ModelThe learned calculation, run here.
Server ≠ model: the machine runs the calculation.◇ Teaching diagram · illustrative shapes and distances, not a model measurement.

Recorded input — the authored target remains outside the initial request.

Show me

Could the input contain more than the visible message?

User message
Write a message to let someone know I will be late.
Implicit system from the official template
You are a helpful AI assistant named SmolLM, trained by Hugging Face
History / tools
Absent in this example.
Inspect the complete input
<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
<|im_start|>user
Write a message to let someone know I will be late.<|im_end|>
<|im_start|>assistant
Roles and exact template
[
  {
    "role": "user",
    "content": "Write a message to let someone know I will be late."
  }
]
{% for message in messages %}{% if loop.first and messages[0]['role'] != 'system' %}{{ '<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
' }}{% endif %}{{'<|im_start|>' + message['role'] + '
' + message['content'] + '<|im_end|>' + '
'}}{% endfor %}{% if add_generation_prompt %}{{ '<|im_start|>assistant
' }}{% endif %}

im_start / im_end delimit roles; the assistant prefix opens the continuation. No message normalization; nothing omitted from the serialization above.

What the application sends and what you see in its message bubble need not be identical.

Go deeper

Show the complete serialized input as well as readable sections. If instructions or history are absent, say absent in this example. Special markers and omitted formatting must be identified.

Other systems run models locally. This diagram does not describe every network or every provider's privacy policy.

Cette scène en français

S03 · The model was trained before this message

Prepared message : Write a message to let someone know I will be late.

What had already been learned before you pressed Send?

A large language model, or LLM, is a computer model trained to process and produce text.

During training, examples, predictions and measured errors are used to adjust numbers called parameters. During this reply, those learned numbers stay fixed.

  1. Before your messageTraining adjusts the parameters.
  2. Trained modelLearned parameters, not stored answer cards.
  3. NowParameters stay fixed during this reply.
Training ≠ generation: the message arrives after learning.◇ Teaching diagram · illustrative shapes and distances, not a model measurement.

Training remains a teaching diagram; the later inference calculations use fixed parameters.

Show me

Does this new message retrain the model?

Before this conversation: training

◇ Teaching diagram — simplified, not measured · □ □ ◇ → □ ◇ ◇

Example → prediction → measured error → parameter adjustment. These marks represent a past change, without invented training numbers.

Present: inference, fixed parameters 🔒

◇ Teaching diagram — simplified, not measured

□ ◇ ◇ → 🔒 · Symbolic marks, not captured values.

During training, examples, predictions and measured errors are used to adjust numbers called parameters. During this reply, those learned numbers stay fixed.

Training changes the model's learned settings. Using them to process an input is called inference.

Go deeper

The training diagram shows one principle, not a capture of this model's training. A prediction is compared with an example's target; an optimization procedure adjusts parameters using a measured error.

Parameters are not one-word drawers or individual memories. Models can nevertheless memorize some training material. Additional training can improve instruction-following; it is not shown here.

Cette scène en français

S04 · Text becomes tokens

Prepared message : Write a message to let someone know I will be late.

Does the model receive a list of whole words?

A tokenizer turns text into units called tokens and their numerical IDs. A token is not necessarily a whole word.

An ID points to an entry in that tokenizer's vocabulary. A larger number does not mean a more important word.

  1. Input textThe full envelope, not just the message bubble.
  2. TokenizerSplits according to its rules.
  3. Ordered tokensA piece, a position and an ID.
Tokenization ≠ understanding. Symbolic tiles here; inspect the exact split below.◇ Teaching diagram · illustrative shapes and distances, not a model measurement.

Verified tokenizer output — exact fragments and IDs, not invented for the story.

Show me

Will every tile match a whole word?

○ Captured trace — raw BPE pieces (Ġ space, Ċ newline). Offsets are Unicode codepoints, sometimes shared; individual decoding can be incomplete.

  1. <|im_start|> · ID 1
  2. system · ID 9690
  3. Ċ · ID 198
  4. You · ID 2683
  5. Ġare · ID 359
  6. Ġa · ID 253
  7. Ġhelpful · ID 5356
  8. ĠAI · ID 5646
  9. Ġassistant · ID 11173
  10. Ġnamed · ID 3365
  11. ĠSm · ID 3511
  12. ol · ID 308
  13. LM · ID 34519
  14. , · ID 28
  15. Ġtrained · ID 7018
  16. Ġby · ID 411
  17. ĠH · ID 407
  18. ugging · ID 19712
  19. ĠFace · ID 8182
  20. <|im_end|> · ID 2
  21. Ċ · ID 198
  22. <|im_start|> · ID 1
  23. user · ID 4093
  24. Ċ · ID 198
  25. Write · ID 19161
  26. Ġa · ID 253
  27. Ġmessage · ID 3714
  28. Ġto · ID 288
  29. Ġlet · ID 1303
  30. Ġsomeone · ID 2206
  31. Ġknow · ID 699
  32. ĠI · ID 339
  33. Ġwill · ID 523
  34. Ġbe · ID 325
  35. Ġlate · ID 3008
  36. . · ID 30
  37. <|im_end|> · ID 2
  38. Ċ · ID 198
  39. <|im_start|> · ID 1
  40. ass · ID 520
  41. istant · ID 9531
  42. Ċ · ID 198

    Inspected position 41 · ID 198 · [198, 199) · Single-token decode: "\n"

<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
<|im_start|>user
Write a message to let someone know I will be late.<|im_end|>
<|im_start|>assistant

Tokenization alone does not understand a sentence. Other tokenizers can segment it differently.

The fragment, its ID and its place in the sequence are different things.

Go deeper

A displayed character can span several tokens, and a token need not decode into a complete Unicode character on its own. Grouped display fragments must list their underlying tokens rather than pretending to be a single token.

Decoding the whole sequence follows this tokenizer's rules. If normalization changes the input, show the original and normalized forms explicitly; do not claim byte-for-byte equality.

Cette scène en français

S05 · An ID looks up a list of numbers

Prepared message : Write a message to let someone know I will be late.

What can the model calculate with an ID?

An ID selects a row in a learned table: a list of numbers called an embedding. This is a starting representation used in the calculations.

The model also accounts for positions in the sequence. The few values shown here are not a complete map of meaning.

  1. Piece → IDThe token keeps its identity.
  2. Learned tableThe ID addresses one row.
  3. RepresentationA list of numbers for calculation.
Looking up a row does not retrain the table. Captured values below, not a map of meaning.◇ Teaching diagram · illustrative shapes and distances, not a model measurement.

Recorded lookup — selected dimensions; layout is not semantic distance.

Show me

Is the token's ID already its full representation?

Display precision: the last digits of numbers may vary after export, without renormalization. Canonical data and SHA256: trace manifest.

Position 41 · Ċ → ID 198 → learned table row. Parameters stay fixed.

□ Documented reduced view of the initial prompt: dimensions 0, 1, 2, 3, 4, 5, 6, 7 / 576. Signed captured values; no human axes or quantitative colour scale.

Position 41 · Ċ → ID 198 → learned table row

Embedding
DimensionEmbedding
0-0.2119140625
1-0.1201171875
20.1728515625
30.056396484375
40.216796875
5-0.0308837890625
60.169921875
7-0.16796875

RoPE handles position through rotations in attention, not psychological axes or a universal addition to this table.

The ID finds a learned row; that list of values can take part in subsequent calculations.

Go deeper

A vector is an ordered list of numbers; a dimension is one numbered place in that list. Each displayed cell must name its dimension and extraction point. Use selected dimensions unless an explicitly documented projection is needed; neither option makes the full space three-dimensional.

The way position enters the calculation depends on the chosen architecture. L02 must document it; the drawing must not assume that position is always added to the embedding.

Cette scène en français

S06 · Context changes the representation

Prepared message : Write a message to let someone know I will be late.

How does the surrounding input affect a token's representation?

The model has successive calculation stages, called layers. Within a layer, attention combines information from allowed positions, never later positions here.

Other learned transformations also change the representation. The original tokens and the learned parameters stay fixed during this reply.

  1. Allowed sourcesThe inspected position and earlier ones. Never the future.
  2. One open layerAttention, then other calculations including the MLP.
  3. Same position, afterA transformed representation, not a new token.
Context changes the representation, not parameters or IDs. Schematic links; inspect recorded weights below.◇ Teaching diagram · illustrative shapes and distances, not a model measurement.

Recorded states and attention — same prefix and position; reduced view, not a complete explanation.

Show me

Can an earlier position read a later token in this model?

Display precision: the last digits of numbers may vary after export, without renormalization. Canonical data and SHA256: trace manifest.

Same inspected position 41 · Ċ · ID 198. Before and after the same layer; text and parameters stay fixed.

Layer 0 · head 0 · zero-based. Block input → RMSNorm, attention, residual, RMSNorm, SiLU MLP, residual → same block output. No new token.

□ Documented reduced view of the initial prompt: dimensions 0, 1, 2, 3, 4, 5, 6, 7 / 576. Signed captured values; no human axes or quantitative colour scale.

Position 41 · Ċ → ID 198 → learned table row

Same position, before/after the layer
DimensionBeforeAfter
0-0.21191406250.3470211923122406
1-0.12011718750.35203033685684204
20.17285156250.6483573317527771
30.0563964843751.6338821649551392
40.2167968750.10558444261550903
5-0.03088378906250.1451607346534729
60.1699218750.2099938243627548
7-0.16796875-0.01370704174041748

RoPE handles position through rotations in attention, not psychological axes or a universal addition to this table.

Allowed attention sources — initial prompt
  1. 0 · <|im_start|> · 0.016806183382868767
  2. 1 · system · 0.03623035550117493
  3. 2 · Ċ · 0.013212859630584717
  4. 3 · You · 0.014478638768196106
  5. 4 · Ġare · 0.0367521271109581
  6. 5 · Ġa · 0.036762870848178864
  7. 6 · Ġhelpful · 0.008883410133421421
  8. 7 · ĠAI · 0.020600130781531337
  9. 8 · Ġassistant · 0.016439218074083328
  10. 9 · Ġnamed · 0.008159841410815716
  11. 10 · ĠSm · 0.009223200380802156
  12. 11 · ol · 0.02037954330444336
  13. 12 · LM · 0.006764519028365612
  14. 13 · , · 0.0029480820521712303
  15. 14 · Ġtrained · 0.03501533716917038
  16. 15 · Ġby · 0.001990683376789093
  17. 16 · ĠH · 0.004558082204312086
  18. 17 · ugging · 0.018426261842250824
  19. 18 · ĠFace · 0.018363045528531075
  20. 19 · <|im_end|> · 0.028487153351306915
  21. 20 · Ċ · 0.013511051423847675
  22. 21 · <|im_start|> · 0.048296812921762466
  23. 22 · user · 0.02644071727991104
  24. 23 · Ċ · 0.00906183198094368
  25. 24 · Write · 0.023222768679261208
  26. 25 · Ġa · 0.05316595360636711
  27. 26 · Ġmessage · 0.038014743477106094
  28. 27 · Ġto · 0.0018556704744696615
  29. 28 · Ġlet · 0.02139618434011936
  30. 29 · Ġsomeone · 0.11133842915296556
  31. 30 · Ġknow · 0.025602929294109344
  32. 31 · ĠI · 0.009219696745276453
  33. 32 · Ġwill · 0.028745058923959732
  34. 33 · Ġbe · 0.020472442731261253
  35. 34 · Ġlate · 0.009819515980780125
  36. 35 · . · 0.00329864164814353
  37. 36 · <|im_end|> · 0.051956940442323685
  38. 37 · Ċ · 0.008122499100863934
  39. 38 · <|im_start|> · 0.05343799665570259
  40. 39 · ass · 0.044940754771232605
  41. 40 · istant · 0.03968193382024765
  42. 41 · Ċ · 0.003915850073099136

× Future positions forbidden: no future position in this input.

The entire allowed row is shown, without renormalization. Parameters and IDs stay fixed; computed representations change.

Context contributes to new calculated values, not a rewrite of the original text or a new training step.

Go deeper

A head is one attention calculation within a layer. Display its identity, the selected row, omitted links and weights without renormalizing a hidden subset as if it were complete.

The feed-forward transformation, often called an MLP, also changes the values. Normalization rescales intermediate values; residual connections carry earlier values alongside a transformation. Their order belongs to the architecture-specific explanation. A single attention map is not a complete causal explanation of a response.

Cette scène en français

S07 · Model scores are not our imposed choice

Prepared message : Write a message to let someone know I will be late.

Who decides the next token in this teaching example?

At the last position of the full input, the model's final representation gives scores for possible next tokens. These scores can be converted to probabilities.

Recorded conditional scores. Imposed choice — not selected by the model. Probability is not truth.

  1. Final representationAt the last position of the prefix.
  2. ○ Calculated maximumThe model’s highest score, not the author of the reply.
  3. ◇ Script targetThe imposed token, even if its score is lower.
Probability ≠ truth. Calculating scores and imposing a choice are separate operations.◇ Teaching diagram · illustrative shapes and distances, not a model measurement.

Recorded conditional scores. Imposed choice — not selected by the model. Probability is not truth.

Show me

Must the imposed token be the model's highest-scoring token?

Display precision: the last digits of numbers may vary after export, without renormalization. Canonical data and SHA256: trace manifest.

Exact context of this choice
<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
<|im_start|>user
Write a message to let someone know I will be late.<|im_end|>
<|im_start|>assistant

Recorded model calculations — given the displayed prefix · Step prefix 0 · Documented reduced view

Final representation at the last position → logits → full-vocabulary softmax.

Probability ≠ truth. These scores concern continuation of the prefix, not the reliability of a claim.

Model maximum (argmax)

" · ID 18 · rank / rang 1

Logit 28.694013595581055 · P 0.2015938808156529 · log(P) -1.601500096342548

The most likely for this prefix. It does not decide which token is appended in our example.

Imposed choice — not selected by the model

Hi · ID 26843 · rank / rang 4

Logit 27.388212203979492 · P 0.05462293180677024 · log(P) -2.9073014879441104

Target already in the top 8; its probability is counted only once.

Imposed choice — not selected by the model · Hi · ID 26843
From representation to probabilities

d0: -2.0552053451538086 d1: -0.027857806533575055 d2: 1.2770594358444214 d3: 0.8531920313835144 d4: -0.3432360589504242 d5: -1.2546347379684448 d6: -0.401529848575592 d7: -1.0341060161590576

Top 8 candidates and imposed target. Fixture values displayed by the browser, scientific notation when needed; no renormalization. Full-vocabulary logsumexp: 30.295513691923603

  • " · ID 18 · rank / rang 1

    Logit 28.694013595581055 · P 0.2015938808156529 · log(P) -1.601500096342548

  • I · ID 57 · rank / rang 2

    Logit 28.587482452392575 · P 0.1812222247894739 · log(P) -1.7080312395310244

  • Hey · ID 22234 · rank / rang 3

    Logit 27.419790267944336 · P 0.05637534147455826 · log(P) -2.8757234239792666

  • Hi · ID 26843 · rank / rang 4

    Logit 27.388212203979492 · P 0.05462293180677024 · log(P) -2.9073014879441104

  • Dear · ID 35097 · rank / rang 5

    Logit 27.24317741394043 · P 0.04724840992530328 · log(P) -3.052336277983173

  • Hello · ID 19556 · rank / rang 6

    Logit 27.06204414367676 · P 0.03942048995464733 · log(P) -3.2334695482468447

  • Sure · ID 34355 · rank / rang 7

    Logit 26.740507125854492 · P 0.02858118724961435 · log(P) -3.5550065660691104

  • Please · ID 10180 · rank / rang 8

    Logit 26.22800064086914 · P 0.01711991195573143 · log(P) -4.067513051054462

Other tokens : 0.3738156220282474 · 49144 omitted candidates.

These tokens are neither in the top 8 nor the imposed target. The target’s probability counts once, even when it appears in both panels. other_probability = mass outside UNION(top8, target).

The maximum breaks ties by lowest ID. The script imposes the target, even when they coincide. Temperature 1, no filters or sampling.

The ordinary loop, unlike our example

In ordinary autoregressive generation, a decoding rule chooses the next token from scores for the current context, extends that context, and repeats. It does not require a complete reply written beforehand.

Inspect ≠ select ≠ append. No token appended yet.

Inspection changes neither the script's imposed choice nor the recorded values.

Go deeper

In ordinary autoregressive generation, a decoding rule chooses the next token from scores for the current context, extends that context, and repeats. It does not require a complete reply written beforehand.

Other mass excludes the union of the top eight candidates and the imposed target. A target outside the top 8 is subtracted once; no softmax over just the visible list.

Cette scène en français

S08 · Add a token, then repeat

Prepared message : Write a message to let someone know I will be late.

What changes after the first token has been chosen?

In ordinary generation, each chosen token joins the context used at the next step.

Recorded calculations for successive prefixes; tokens imposed by the script. End of prepared path — not a model-selected stop.

  1. Current prefixThe input and tokens already appended.
  2. ◇ Imposed choiceThe script fixes the token; its scores are recorded.
  3. Extended prefixAppending changes the context for the next calculation.
Training ≠ generation: parameters stay fixed. Here we replay a guided path, not autonomous generation.◇ Teaching diagram · illustrative shapes and distances, not a model measurement.

Recorded calculations for successive prefixes; tokens imposed by the script. End of prepared path — not a model-selected stop.

Show me

Can the next step keep using the old context?

Display precision: the last digits of numbers may vary after export, without renormalization. Canonical data and SHA256: trace manifest.

Imposed tokens appended : 16 / 16.

Hi, I will be a little late. Sorry to keep you waiting!
Prefix that produced the last appended token · ID 17
<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
<|im_start|>user
Write a message to let someone know I will be late.<|im_end|>
<|im_start|>assistant
Hi, I will be a little late. Sorry to keep you waiting
Final prefix — no next choice captured · 42 + 16 tokens
<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
<|im_start|>user
Write a message to let someone know I will be late.<|im_end|>
<|im_start|>assistant
Hi, I will be a little late. Sorry to keep you waiting!

1, 9690, 198, 2683, 359, 253, 5356, 5646, 11173, 3365, 3511, 308, 34519, 28, 7018, 411, 407, 19712, 8182, 2, 198, 1, 4093, 198, 19161, 253, 3714, 288, 1303, 2206, 699, 339, 523, 325, 3008, 30, 2, 198, 1, 520, 9531, 198, 26843, 28, 339, 523, 325, 253, 1838, 3008, 30, 25442, 541, 288, 1446, 346, 7003, 17

End of prepared path — not a model-selected stop: prepared_sequence_end · 16 tokens. No EOS appended, no invented terminal distribution.

Exact ledger of every step
g=0 → 1 · ID 26843
<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
<|im_start|>user
Write a message to let someone know I will be late.<|im_end|>
<|im_start|>assistant

1, 9690, 198, 2683, 359, 253, 5356, 5646, 11173, 3365, 3511, 308, 34519, 28, 7018, 411, 407, 19712, 8182, 2, 198, 1, 4093, 198, 19161, 253, 3714, 288, 1303, 2206, 699, 339, 523, 325, 3008, 30, 2, 198, 1, 520, 9531, 198

Recorded model calculations — given the displayed prefix · Step prefix 0 · Documented reduced view

Final representation at the last position → logits → full-vocabulary softmax.

Probability ≠ truth. These scores concern continuation of the prefix, not the reliability of a claim.

Model maximum (argmax)

" · ID 18 · rank / rang 1

Logit 28.694013595581055 · P 0.2015938808156529 · log(P) -1.601500096342548

The most likely for this prefix. It does not decide which token is appended in our example.

Imposed choice — not selected by the model

Hi · ID 26843 · rank / rang 4

Logit 27.388212203979492 · P 0.05462293180677024 · log(P) -2.9073014879441104

Target already in the top 8; its probability is counted only once.

Imposed choice — not selected by the model · Hi · ID 26843
From representation to probabilities

d0: -2.0552053451538086 d1: -0.027857806533575055 d2: 1.2770594358444214 d3: 0.8531920313835144 d4: -0.3432360589504242 d5: -1.2546347379684448 d6: -0.401529848575592 d7: -1.0341060161590576

Top 8 candidates and imposed target. Fixture values displayed by the browser, scientific notation when needed; no renormalization. Full-vocabulary logsumexp: 30.295513691923603

  • " · ID 18 · rank / rang 1

    Logit 28.694013595581055 · P 0.2015938808156529 · log(P) -1.601500096342548

  • I · ID 57 · rank / rang 2

    Logit 28.587482452392575 · P 0.1812222247894739 · log(P) -1.7080312395310244

  • Hey · ID 22234 · rank / rang 3

    Logit 27.419790267944336 · P 0.05637534147455826 · log(P) -2.8757234239792666

  • Hi · ID 26843 · rank / rang 4

    Logit 27.388212203979492 · P 0.05462293180677024 · log(P) -2.9073014879441104

  • Dear · ID 35097 · rank / rang 5

    Logit 27.24317741394043 · P 0.04724840992530328 · log(P) -3.052336277983173

  • Hello · ID 19556 · rank / rang 6

    Logit 27.06204414367676 · P 0.03942048995464733 · log(P) -3.2334695482468447

  • Sure · ID 34355 · rank / rang 7

    Logit 26.740507125854492 · P 0.02858118724961435 · log(P) -3.5550065660691104

  • Please · ID 10180 · rank / rang 8

    Logit 26.22800064086914 · P 0.01711991195573143 · log(P) -4.067513051054462

Other tokens : 0.3738156220282474 · 49144 omitted candidates.

These tokens are neither in the top 8 nor the imposed target. The target’s probability counts once, even when it appears in both panels. other_probability = mass outside UNION(top8, target).

The maximum breaks ties by lowest ID. The script imposes the target, even when they coincide. Temperature 1, no filters or sampling.

The ordinary loop, unlike our example

In ordinary autoregressive generation, a decoding rule chooses the next token from scores for the current context, extends that context, and repeats. It does not require a complete reply written beforehand.

Hi
g=1 → 2 · ID 28
<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
<|im_start|>user
Write a message to let someone know I will be late.<|im_end|>
<|im_start|>assistant
Hi

1, 9690, 198, 2683, 359, 253, 5356, 5646, 11173, 3365, 3511, 308, 34519, 28, 7018, 411, 407, 19712, 8182, 2, 198, 1, 4093, 198, 19161, 253, 3714, 288, 1303, 2206, 699, 339, 523, 325, 3008, 30, 2, 198, 1, 520, 9531, 198, 26843

Recorded model calculations — given the displayed prefix · Step prefix 1 · Documented reduced view

Final representation at the last position → logits → full-vocabulary softmax.

Probability ≠ truth. These scores concern continuation of the prefix, not the reliability of a claim.

Model maximum (argmax)

Ġ[ · ID 933 · rank / rang 1

Logit 7.943227291107178 · P 0.7508391956959775 · log(P) -0.28656377039005143

The most likely for this prefix. It does not decide which token is appended in our example.

Imposed choice — not selected by the model

, · ID 28 · rank / rang 6

Logit 3.8363325595855713 · P 0.012357915353324738 · log(P) -4.393458501911658

Target already in the top 8; its probability is counted only once.

Imposed choice — not selected by the model · , · ID 28
From representation to probabilities

d0: -1.3738007545471191 d1: -0.5202157497406006 d2: 1.876486659049988 d3: 0.9552001953125 d4: -0.08893979340791702 d5: -2.5445735454559326 d6: -1.5463964939117432 d7: -1.7778035402297974

Top 8 candidates and imposed target. Fixture values displayed by the browser, scientific notation when needed; no renormalization. Full-vocabulary logsumexp: 8.22979106149723

  • Ġ[ · ID 933 · rank / rang 1

    Logit 7.943227291107178 · P 0.7508391956959775 · log(P) -0.28656377039005143

  • ĠAlex · ID 5325 · rank / rang 2

    Logit 4.963054656982422 · P 0.038130667303635395 · log(P) -3.2667364045148073

  • ĠEmily · ID 15487 · rank / rang 3

    Logit 4.271503448486328 · P 0.0190957856855478 · log(P) -3.958287613010901

  • ĠJohn · ID 2322 · rank / rang 4

    Logit 4.26744270324707 · P 0.01901839979327099 · log(P) -3.962348358250159

  • ĠSarah · ID 8165 · rank / rang 5

    Logit 3.86134672164917 · P 0.01267093691526863 · log(P) -4.368444339848059

  • , · ID 28 · rank / rang 6

    Logit 3.8363325595855713 · P 0.012357915353324738 · log(P) -4.393458501911658

  • Ġthere · ID 665 · rank / rang 7

    Logit 3.5609374046325684 · P 0.009383019524256448 · log(P) -4.668853656864661

  • ĠLi · ID 9679 · rank / rang 8

    Logit 3.2445428371429443 · P 0.006838080438675774 · log(P) -4.985248224354285

Other tokens : 0.1316659992900421 · 49144 omitted candidates.

These tokens are neither in the top 8 nor the imposed target. The target’s probability counts once, even when it appears in both panels. other_probability = mass outside UNION(top8, target).

The maximum breaks ties by lowest ID. The script imposes the target, even when they coincide. Temperature 1, no filters or sampling.

The ordinary loop, unlike our example

In ordinary autoregressive generation, a decoding rule chooses the next token from scores for the current context, extends that context, and repeats. It does not require a complete reply written beforehand.

Hi,
g=2 → 3 · ID 339
<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
<|im_start|>user
Write a message to let someone know I will be late.<|im_end|>
<|im_start|>assistant
Hi,

1, 9690, 198, 2683, 359, 253, 5356, 5646, 11173, 3365, 3511, 308, 34519, 28, 7018, 411, 407, 19712, 8182, 2, 198, 1, 4093, 198, 19161, 253, 3714, 288, 1303, 2206, 699, 339, 523, 325, 3008, 30, 2, 198, 1, 520, 9531, 198, 26843, 28

Recorded model calculations — given the displayed prefix · Step prefix 2 · Documented reduced view

Final representation at the last position → logits → full-vocabulary softmax.

Probability ≠ truth. These scores concern continuation of the prefix, not the reliability of a claim.

Model maximum (argmax)

ĠI · ID 339 · rank / rang 1

Logit 12.261602401733398 · P 0.5416538339823407 · log(P) -0.6131281642765725

The most likely for this prefix. It does not decide which token is appended in our example.

Imposed choice — not selected by the model

ĠI · ID 339 · rank / rang 1

Logit 12.261602401733398 · P 0.5416538339823407 · log(P) -0.6131281642765725

Target already in the top 8; its probability is counted only once.

Imposed choice — not selected by the model · ĠI · ID 339
From representation to probabilities

d0: -1.3890047073364258 d1: 0.04630286619067192 d2: 1.4595191478729248 d3: 0.8313894271850586 d4: 0.3924960494041443 d5: -1.2829383611679075 d6: -0.03668787330389023 d7: -3.1501736640930176

Top 8 candidates and imposed target. Fixture values displayed by the browser, scientific notation when needed; no renormalization. Full-vocabulary logsumexp: 12.874730566009973

  • ĠI · ID 339 · rank / rang 1

    Logit 12.261602401733398 · P 0.5416538339823407 · log(P) -0.6131281642765725

  • Ġit · ID 357 · rank / rang 2

    Logit 10.191718101501465 · P 0.06835692087697401 · log(P) -2.683012464508506

  • Ġthanks · ID 6954 · rank / rang 3

    Logit 10.008417129516602 · P 0.056908336537826 · log(P) -2.8663134364933693

  • Ġhow · ID 638 · rank / rang 4

    Logit 9.530943870544434 · P 0.03530302250047117 · log(P) -3.3437866954655373

  • Ċ · ID 198 · rank / rang 5

    Logit 8.772052764892578 · P 0.01652835643822013 · log(P) -4.102677801117393

  • Ġsorry · ID 22657 · rank / rang 6

    Logit 8.762948036193848 · P 0.016378553230636035 · log(P) -4.111782529816123

  • Ġcan · ID 416 · rank / rang 7

    Logit 8.560697555541992 · P 0.01337948105089643 · log(P) -4.314033010467979

  • Ġthank · ID 9984 · rank / rang 8

    Logit 8.480876922607422 · P 0.012353033192863324 · log(P) -4.393853643402549

Other tokens : 0.23913846218977125 · 49144 omitted candidates.

These tokens are neither in the top 8 nor the imposed target. The target’s probability counts once, even when it appears in both panels. other_probability = mass outside UNION(top8, target).

The maximum breaks ties by lowest ID. The script imposes the target, even when they coincide. Temperature 1, no filters or sampling.

The ordinary loop, unlike our example

In ordinary autoregressive generation, a decoding rule chooses the next token from scores for the current context, extends that context, and repeats. It does not require a complete reply written beforehand.

Hi, I
g=3 → 4 · ID 523
<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
<|im_start|>user
Write a message to let someone know I will be late.<|im_end|>
<|im_start|>assistant
Hi, I

1, 9690, 198, 2683, 359, 253, 5356, 5646, 11173, 3365, 3511, 308, 34519, 28, 7018, 411, 407, 19712, 8182, 2, 198, 1, 4093, 198, 19161, 253, 3714, 288, 1303, 2206, 699, 339, 523, 325, 3008, 30, 2, 198, 1, 520, 9531, 198, 26843, 28, 339

Recorded model calculations — given the displayed prefix · Step prefix 3 · Documented reduced view

Final representation at the last position → logits → full-vocabulary softmax.

Probability ≠ truth. These scores concern continuation of the prefix, not the reliability of a claim.

Model maximum (argmax)

'm · ID 5248 · rank / rang 1

Logit 29.538902282714844 · P 0.6647063412775203 · log(P) -0.4084099279206512

The most likely for this prefix. It does not decide which token is appended in our example.

Imposed choice — not selected by the model

Ġwill · ID 523 · rank / rang 11

Logit 25.369293212890625 · P 0.010275231404390506 · log(P) -4.57801899774487

Target outside top 8 — not promoted in the ranking.

Imposed choice — not selected by the model · Ġwill · ID 523
From representation to probabilities

d0: 1.0284905433654783 d1: -0.5418347716331482 d2: -0.4579586684703827 d3: 0.5164719820022583 d4: 1.357612371444702 d5: -1.0663337707519531 d6: 2.732614517211914 d7: -2.6122002601623535

Top 8 candidates and imposed target. Fixture values displayed by the browser, scientific notation when needed; no renormalization. Full-vocabulary logsumexp: 29.947312210635495

  • 'm · ID 5248 · rank / rang 1

    Logit 29.538902282714844 · P 0.6647063412775203 · log(P) -0.4084099279206512

  • Ġhope · ID 3826 · rank / rang 2

    Logit 27.380939483642575 · P 0.07681366484450634 · log(P) -2.566372726992917

  • Ġam · ID 744 · rank / rang 3

    Logit 26.44110870361328 · P 0.03001063359304285 · log(P) -3.5062035070222137

  • 'll · ID 3060 · rank / rang 4

    Logit 26.440147399902344 · P 0.029981798121645684 · log(P) -3.5071648107331512

  • Ġwas · ID 436 · rank / rang 5

    Logit 26.164474487304688 · P 0.022758018746148117 · log(P) -3.782837723330807

  • 've · ID 3543 · rank / rang 6

    Logit 26.095613479614254 · P 0.02124361857591764 · log(P) -3.851698731021237

  • Ġjust · ID 915 · rank / rang 7

    Logit 25.748607635498047 · P 0.015015015051184615 · log(P) -4.198704575137448

  • Ġwanted · ID 4146 · rank / rang 8

    Logit 25.705127716064453 · P 0.01437615288856432 · log(P) -4.242184494571042

Other tokens : 0.11481952549707986 · 49143 omitted candidates.

These tokens are neither in the top 8 nor the imposed target. The target’s probability counts once, even when it appears in both panels. other_probability = mass outside UNION(top8, target).

The maximum breaks ties by lowest ID. The script imposes the target, even when they coincide. Temperature 1, no filters or sampling.

The ordinary loop, unlike our example

In ordinary autoregressive generation, a decoding rule chooses the next token from scores for the current context, extends that context, and repeats. It does not require a complete reply written beforehand.

Hi, I will
g=4 → 5 · ID 325
<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
<|im_start|>user
Write a message to let someone know I will be late.<|im_end|>
<|im_start|>assistant
Hi, I will

1, 9690, 198, 2683, 359, 253, 5356, 5646, 11173, 3365, 3511, 308, 34519, 28, 7018, 411, 407, 19712, 8182, 2, 198, 1, 4093, 198, 19161, 253, 3714, 288, 1303, 2206, 699, 339, 523, 325, 3008, 30, 2, 198, 1, 520, 9531, 198, 26843, 28, 339, 523

Recorded model calculations — given the displayed prefix · Step prefix 4 · Documented reduced view

Final representation at the last position → logits → full-vocabulary softmax.

Probability ≠ truth. These scores concern continuation of the prefix, not the reliability of a claim.

Model maximum (argmax)

Ġbe · ID 325 · rank / rang 1

Logit 27.738601684570312 · P 0.9486104014679388 · log(P) -0.05275710052609739

The most likely for this prefix. It does not decide which token is appended in our example.

Imposed choice — not selected by the model

Ġbe · ID 325 · rank / rang 1

Logit 27.738601684570312 · P 0.9486104014679388 · log(P) -0.05275710052609739

Target already in the top 8; its probability is counted only once.

Imposed choice — not selected by the model · Ġbe · ID 325
From representation to probabilities

d0: -0.7276690602302551 d1: -2.4926223754882812 d2: -0.4160797894001007 d3: -0.05029306933283806 d4: 0.969337522983551 d5: -1.3014726638793943 d6: 0.22555148601531985 d7: -2.047252893447876

Top 8 candidates and imposed target. Fixture values displayed by the browser, scientific notation when needed; no renormalization. Full-vocabulary logsumexp: 27.79135878509641

  • Ġbe · ID 325 · rank / rang 1

    Logit 27.738601684570312 · P 0.9486104014679388 · log(P) -0.05275710052609739

  • Ġneed · ID 737 · rank / rang 2

    Logit 23.395551681518555 · P 0.012328925588765326 · log(P) -4.395807103577855

  • Ġhave · ID 457 · rank / rang 3

    Logit 23.09413719177246 · P 0.009120582618545262 · log(P) -4.697221593323949

  • Ġmiss · ID 3956 · rank / rang 4

    Logit 22.54572296142578 · P 0.005270469539751897 · log(P) -5.245635823670629

  • Ġnot · ID 441 · rank / rang 5

    Logit 22.185312271118164 · P 0.0036755719933553375 · log(P) -5.606046513978246

  • Ġprobably · ID 3105 · rank / rang 6

    Logit 21.73445701599121 · P 0.002341644616057088 · log(P) -6.056901769105199

  • Ġtry · ID 1576 · rank / rang 7

    Logit 21.62902069091797 · P 0.002107320393668632 · log(P) -6.162338094178441

  • Ġget · ID 820 · rank / rang 8

    Logit 21.475494384765625 · P 0.0018074027656234824 · log(P) -6.315864400330785

Other tokens : 0.014737681016295376 · 49144 omitted candidates.

These tokens are neither in the top 8 nor the imposed target. The target’s probability counts once, even when it appears in both panels. other_probability = mass outside UNION(top8, target).

The maximum breaks ties by lowest ID. The script imposes the target, even when they coincide. Temperature 1, no filters or sampling.

The ordinary loop, unlike our example

In ordinary autoregressive generation, a decoding rule chooses the next token from scores for the current context, extends that context, and repeats. It does not require a complete reply written beforehand.

Hi, I will be
g=5 → 6 · ID 253
<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
<|im_start|>user
Write a message to let someone know I will be late.<|im_end|>
<|im_start|>assistant
Hi, I will be

1, 9690, 198, 2683, 359, 253, 5356, 5646, 11173, 3365, 3511, 308, 34519, 28, 7018, 411, 407, 19712, 8182, 2, 198, 1, 4093, 198, 19161, 253, 3714, 288, 1303, 2206, 699, 339, 523, 325, 3008, 30, 2, 198, 1, 520, 9531, 198, 26843, 28, 339, 523, 325

Recorded model calculations — given the displayed prefix · Step prefix 5 · Documented reduced view

Final representation at the last position → logits → full-vocabulary softmax.

Probability ≠ truth. These scores concern continuation of the prefix, not the reliability of a claim.

Model maximum (argmax)

Ġarriving · ID 18563 · rank / rang 1

Logit 22.357425689697266 · P 0.49350416399246727 · log(P) -0.706223982469492

The most likely for this prefix. It does not decide which token is appended in our example.

Imposed choice — not selected by the model

Ġa · ID 253 · rank / rang 17

Logit 17.270858764648438 · P 0.003049459529467995 · log(P) -5.79279090751832

Target outside top 8 — not promoted in the ranking.

Imposed choice — not selected by the model · Ġa · ID 253
From representation to probabilities

d0: -0.6928384304046631 d1: -0.9491083025932312 d2: -0.5656501650810242 d3: 0.4913584887981415 d4: 1.20679771900177 d5: -0.6256847977638245 d6: 0.007863008417189121 d7: -2.1528191566467285

Top 8 candidates and imposed target. Fixture values displayed by the browser, scientific notation when needed; no renormalization. Full-vocabulary logsumexp: 23.06364967216676

  • Ġarriving · ID 18563 · rank / rang 1

    Logit 22.357425689697266 · P 0.49350416399246727 · log(P) -0.706223982469492

  • Ġlate · ID 3008 · rank / rang 2

    Logit 21.149677276611328 · P 0.14749331952919614 · log(P) -1.9139723955554295

  • Ġat · ID 418 · rank / rang 3

    Logit 20.44598960876465 · P 0.07297341637502132 · log(P) -2.617660063402109

  • Ġin · ID 281 · rank / rang 4

    Logit 19.95261001586914 · P 0.044554609772080714 · log(P) -3.111039656297617

  • Ġon · ID 335 · rank / rang 5

    Logit 19.290359497070312 · P 0.02297634259814696 · log(P) -3.773290175096445

  • Ġgoing · ID 2045 · rank / rang 6

    Logit 18.950429916381836 · P 0.016355030456421685 · log(P) -4.113219755784922

  • Ġcoming · ID 4167 · rank / rang 7

    Logit 18.836090087890625 · P 0.014587947800722006 · log(P) -4.227559584276133

  • Ġhere · ID 1535 · rank / rang 8

    Logit 18.456806182861328 · P 0.00998328095874125 · log(P) -4.6068434893054295

Other tokens : 0.17452242898773507 · 49143 omitted candidates.

These tokens are neither in the top 8 nor the imposed target. The target’s probability counts once, even when it appears in both panels. other_probability = mass outside UNION(top8, target).

The maximum breaks ties by lowest ID. The script imposes the target, even when they coincide. Temperature 1, no filters or sampling.

The ordinary loop, unlike our example

In ordinary autoregressive generation, a decoding rule chooses the next token from scores for the current context, extends that context, and repeats. It does not require a complete reply written beforehand.

Hi, I will be a
g=6 → 7 · ID 1838
<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
<|im_start|>user
Write a message to let someone know I will be late.<|im_end|>
<|im_start|>assistant
Hi, I will be a

1, 9690, 198, 2683, 359, 253, 5356, 5646, 11173, 3365, 3511, 308, 34519, 28, 7018, 411, 407, 19712, 8182, 2, 198, 1, 4093, 198, 19161, 253, 3714, 288, 1303, 2206, 699, 339, 523, 325, 3008, 30, 2, 198, 1, 520, 9531, 198, 26843, 28, 339, 523, 325, 253

Recorded model calculations — given the displayed prefix · Step prefix 6 · Documented reduced view

Final representation at the last position → logits → full-vocabulary softmax.

Probability ≠ truth. These scores concern continuation of the prefix, not the reliability of a claim.

Model maximum (argmax)

Ġbit · ID 3241 · rank / rang 1

Logit 24.410703659057617 · P 0.3844536556696628 · log(P) -0.9559320287192355

The most likely for this prefix. It does not decide which token is appended in our example.

Imposed choice — not selected by the model

Ġlittle · ID 1838 · rank / rang 2

Logit 24.38994407653809 · P 0.3765548301116357 · log(P) -0.9766916112387668

Target already in the top 8; its probability is counted only once.

Imposed choice — not selected by the model · Ġlittle · ID 1838
From representation to probabilities

d0: -2.269467353820801 d1: -0.06955106556415558 d2: -0.49502983689308167 d3: 1.1438674926757812 d4: 0.37673303484916687 d5: -0.2138453871011734 d6: -2.409039258956909 d7: -2.5745697021484375

Top 8 candidates and imposed target. Fixture values displayed by the browser, scientific notation when needed; no renormalization. Full-vocabulary logsumexp: 25.366635687776853

  • Ġbit · ID 3241 · rank / rang 1

    Logit 24.410703659057617 · P 0.3844536556696628 · log(P) -0.9559320287192355

  • Ġlittle · ID 1838 · rank / rang 2

    Logit 24.38994407653809 · P 0.3765548301116357 · log(P) -0.9766916112387668

  • Ġlot · ID 2341 · rank / rang 3

    Logit 22.784648895263672 · P 0.07562360625301334 · log(P) -2.581986792513181

  • Ġlate · ID 3008 · rank / rang 4

    Logit 22.3887939453125 · P 0.05090257626492618 · log(P) -2.977841742464353

  • Ġfew · ID 1443 · rank / rang 5

    Logit 20.868349075317383 · P 0.011128046870138878 · log(P) -4.49828661245947

  • - · ID 29 · rank / rang 6

    Logit 20.341785430908203 · P 0.006572570618550343 · log(P) -5.02485025686865

  • arr · ID 2110 · rank / rang 7

    Logit 19.67495346069336 · P 0.0033739125058068658 · log(P) -5.6916822270834935

  • Ġday · ID 1194 · rank / rang 8

    Logit 19.67462921142578 · P 0.003372818694491328 · log(P) -5.692006476351072

Other tokens : 0.08801798301177345 · 49144 omitted candidates.

These tokens are neither in the top 8 nor the imposed target. The target’s probability counts once, even when it appears in both panels. other_probability = mass outside UNION(top8, target).

The maximum breaks ties by lowest ID. The script imposes the target, even when they coincide. Temperature 1, no filters or sampling.

The ordinary loop, unlike our example

In ordinary autoregressive generation, a decoding rule chooses the next token from scores for the current context, extends that context, and repeats. It does not require a complete reply written beforehand.

Hi, I will be a little
g=7 → 8 · ID 3008
<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
<|im_start|>user
Write a message to let someone know I will be late.<|im_end|>
<|im_start|>assistant
Hi, I will be a little

1, 9690, 198, 2683, 359, 253, 5356, 5646, 11173, 3365, 3511, 308, 34519, 28, 7018, 411, 407, 19712, 8182, 2, 198, 1, 4093, 198, 19161, 253, 3714, 288, 1303, 2206, 699, 339, 523, 325, 3008, 30, 2, 198, 1, 520, 9531, 198, 26843, 28, 339, 523, 325, 253, 1838

Recorded model calculations — given the displayed prefix · Step prefix 7 · Documented reduced view

Final representation at the last position → logits → full-vocabulary softmax.

Probability ≠ truth. These scores concern continuation of the prefix, not the reliability of a claim.

Model maximum (argmax)

Ġbit · ID 3241 · rank / rang 1

Logit 24.832923889160156 · P 0.377625762409583 · log(P) -0.9738516203178342

The most likely for this prefix. It does not decide which token is appended in our example.

Imposed choice — not selected by the model

Ġlate · ID 3008 · rank / rang 2

Logit 24.48090171813965 · P 0.2655708042078022 · log(P) -1.325873791338342

Target already in the top 8; its probability is counted only once.

Imposed choice — not selected by the model · Ġlate · ID 3008
From representation to probabilities

d0: -2.0251805782318115 d1: -1.4032589197158811 d2: 0.05142971500754357 d3: 1.056714415550232 d4: 0.015530036762356758 d5: 0.029730547219514847 d6: 0.8244141340255737 d7: -1.4314391613006592

Top 8 candidates and imposed target. Fixture values displayed by the browser, scientific notation when needed; no renormalization. Full-vocabulary logsumexp: 25.80677550947799

  • Ġbit · ID 3241 · rank / rang 1

    Logit 24.832923889160156 · P 0.377625762409583 · log(P) -0.9738516203178342

  • Ġlate · ID 3008 · rank / rang 2

    Logit 24.48090171813965 · P 0.2655708042078022 · log(P) -1.325873791338342

  • Ġearly · ID 1584 · rank / rang 3

    Logit 23.934722900390625 · P 0.15380763064149974 · log(P) -1.8720526090873657

  • Ġearlier · ID 3685 · rank / rang 4

    Logit 22.39000701904297 · P 0.03281831647187173 · log(P) -3.4167684904350217

  • Ġlater · ID 1916 · rank / rang 5

    Logit 22.140840530395508 · P 0.025580243107053553 · log(P) -3.6659349790824822

  • Ġbefore · ID 1092 · rank / rang 6

    Logit 21.63579559326172 · P 0.01543712562011995 · log(P) -4.170979916216272

  • Ġover · ID 690 · rank / rang 7

    Logit 21.36518096923828 · P 0.011777144409081643 · log(P) -4.441594540239709

  • Ġmore · ID 540 · rank / rang 8

    Logit 21.344587326049805 · P 0.011537090376698696 · log(P) -4.462188183428186

Other tokens : 0.10584588275629098 · 49144 omitted candidates.

These tokens are neither in the top 8 nor the imposed target. The target’s probability counts once, even when it appears in both panels. other_probability = mass outside UNION(top8, target).

The maximum breaks ties by lowest ID. The script imposes the target, even when they coincide. Temperature 1, no filters or sampling.

The ordinary loop, unlike our example

In ordinary autoregressive generation, a decoding rule chooses the next token from scores for the current context, extends that context, and repeats. It does not require a complete reply written beforehand.

Hi, I will be a little late
g=8 → 9 · ID 30
<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
<|im_start|>user
Write a message to let someone know I will be late.<|im_end|>
<|im_start|>assistant
Hi, I will be a little late

1, 9690, 198, 2683, 359, 253, 5356, 5646, 11173, 3365, 3511, 308, 34519, 28, 7018, 411, 407, 19712, 8182, 2, 198, 1, 4093, 198, 19161, 253, 3714, 288, 1303, 2206, 699, 339, 523, 325, 3008, 30, 2, 198, 1, 520, 9531, 198, 26843, 28, 339, 523, 325, 253, 1838, 3008

Recorded model calculations — given the displayed prefix · Step prefix 8 · Documented reduced view

Final representation at the last position → logits → full-vocabulary softmax.

Probability ≠ truth. These scores concern continuation of the prefix, not the reliability of a claim.

Model maximum (argmax)

. · ID 30 · rank / rang 1

Logit 19.96644592285156 · P 0.314481590195263 · log(P) -1.156829741295006

The most likely for this prefix. It does not decide which token is appended in our example.

Imposed choice — not selected by the model

. · ID 30 · rank / rang 1

Logit 19.96644592285156 · P 0.314481590195263 · log(P) -1.156829741295006

Target already in the top 8; its probability is counted only once.

Imposed choice — not selected by the model · . · ID 30
From representation to probabilities

d0: -1.2886192798614502 d1: -0.30560430884361267 d2: -1.0620988607406616 d3: 1.7276201248168943 d4: 0.29113897681236267 d5: -1.7176393270492554 d6: -1.6500927209854126 d7: -2.166079044342041

Top 8 candidates and imposed target. Fixture values displayed by the browser, scientific notation when needed; no renormalization. Full-vocabulary logsumexp: 21.12327566414657

  • . · ID 30 · rank / rang 1

    Logit 19.96644592285156 · P 0.314481590195263 · log(P) -1.156829741295006

  • Ġfor · ID 327 · rank / rang 2

    Logit 18.985572814941406 · P 0.11792542460230356 · log(P) -2.137702849205162

  • Ġto · ID 288 · rank / rang 3

    Logit 18.731313705444336 · P 0.09145008646080596 · log(P) -2.3919619587022325

  • Ġand · ID 284 · rank / rang 4

    Logit 18.48525047302246 · P 0.0715023336663849 · log(P) -2.6380251911241075

  • , · ID 28 · rank / rang 5

    Logit 18.44091796875 · P 0.06840169353362814 · log(P) -2.6823576953965684

  • Ġtoday · ID 1834 · rank / rang 6

    Logit 17.808347702026367 · P 0.03633666609944325 · log(P) -3.3149279621202012

  • Ġthis · ID 451 · rank / rang 7

    Logit 17.516998291015625 · P 0.02715273847274255 · log(P) -3.606277373130943

  • Ġtonight · ID 34451 · rank / rang 8

    Logit 17.304285049438477 · P 0.021949945668392624 · log(P) -3.818990614708092

Other tokens : 0.2507995213010349 · 49144 omitted candidates.

These tokens are neither in the top 8 nor the imposed target. The target’s probability counts once, even when it appears in both panels. other_probability = mass outside UNION(top8, target).

The maximum breaks ties by lowest ID. The script imposes the target, even when they coincide. Temperature 1, no filters or sampling.

The ordinary loop, unlike our example

In ordinary autoregressive generation, a decoding rule chooses the next token from scores for the current context, extends that context, and repeats. It does not require a complete reply written beforehand.

Hi, I will be a little late.
g=9 → 10 · ID 25442
<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
<|im_start|>user
Write a message to let someone know I will be late.<|im_end|>
<|im_start|>assistant
Hi, I will be a little late.

1, 9690, 198, 2683, 359, 253, 5356, 5646, 11173, 3365, 3511, 308, 34519, 28, 7018, 411, 407, 19712, 8182, 2, 198, 1, 4093, 198, 19161, 253, 3714, 288, 1303, 2206, 699, 339, 523, 325, 3008, 30, 2, 198, 1, 520, 9531, 198, 26843, 28, 339, 523, 325, 253, 1838, 3008, 30

Recorded model calculations — given the displayed prefix · Step prefix 9 · Documented reduced view

Final representation at the last position → logits → full-vocabulary softmax.

Probability ≠ truth. These scores concern continuation of the prefix, not the reliability of a claim.

Model maximum (argmax)

ĠI · ID 339 · rank / rang 1

Logit 15.978963851928713 · P 0.24033545664473255 · log(P) -1.425719595544134

The most likely for this prefix. It does not decide which token is appended in our example.

Imposed choice — not selected by the model

ĠSor · ID 25442 · rank / rang 8

Logit 13.81415557861328 · P 0.02758376596844217 · log(P) -3.5905278688595637

Target already in the top 8; its probability is counted only once.

Imposed choice — not selected by the model · ĠSor · ID 25442
From representation to probabilities

d0: -1.1852322816848757 d1: -0.6666378974914551 d2: 0.8172832131385803 d3: 1.3176355361938477 d4: 0.6936988234519958 d5: -2.0851871967315674 d6: -0.9130510687828064 d7: -1.7433034181594849

Top 8 candidates and imposed target. Fixture values displayed by the browser, scientific notation when needed; no renormalization. Full-vocabulary logsumexp: 17.404683447472845

  • ĠI · ID 339 · rank / rang 1

    Logit 15.978963851928713 · P 0.24033545664473255 · log(P) -1.425719595544134

  • ĠCan · ID 1978 · rank / rang 2

    Logit 15.119123458862305 · P 0.10171708574520136 · log(P) -2.28555998861054

  • ĠPlease · ID 8765 · rank / rang 3

    Logit 15.051177024841309 · P 0.09503534322406494 · log(P) -2.3535064226315363

  • <|im_end|> · ID 2 · rank / rang 4

    Logit 14.410785675048828 · P 0.05009180924283071 · log(P) -2.993897772424017

  • ĠIt · ID 657 · rank / rang 5

    Logit 14.340883255004885 · P 0.04670985108876475 · log(P) -3.063800192467962

  • ĠHow · ID 1073 · rank / rang 6

    Logit 14.08203125 · P 0.03605707434494237 · log(P) -3.322652197472845

  • ĠYou · ID 1206 · rank / rang 7

    Logit 13.96531105041504 · P 0.032084815534940914 · log(P) -3.439372397057806

  • ĠSor · ID 25442 · rank / rang 8

    Logit 13.81415557861328 · P 0.02758376596844217 · log(P) -3.5905278688595637

Other tokens : 0.3703847982060817 · 49144 omitted candidates.

These tokens are neither in the top 8 nor the imposed target. The target’s probability counts once, even when it appears in both panels. other_probability = mass outside UNION(top8, target).

The maximum breaks ties by lowest ID. The script imposes the target, even when they coincide. Temperature 1, no filters or sampling.

The ordinary loop, unlike our example

In ordinary autoregressive generation, a decoding rule chooses the next token from scores for the current context, extends that context, and repeats. It does not require a complete reply written beforehand.

Hi, I will be a little late. Sor
g=10 → 11 · ID 541
<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
<|im_start|>user
Write a message to let someone know I will be late.<|im_end|>
<|im_start|>assistant
Hi, I will be a little late. Sor

1, 9690, 198, 2683, 359, 253, 5356, 5646, 11173, 3365, 3511, 308, 34519, 28, 7018, 411, 407, 19712, 8182, 2, 198, 1, 4093, 198, 19161, 253, 3714, 288, 1303, 2206, 699, 339, 523, 325, 3008, 30, 2, 198, 1, 520, 9531, 198, 26843, 28, 339, 523, 325, 253, 1838, 3008, 30, 25442

Recorded model calculations — given the displayed prefix · Step prefix 10 · Documented reduced view

Final representation at the last position → logits → full-vocabulary softmax.

Probability ≠ truth. These scores concern continuation of the prefix, not the reliability of a claim.

Model maximum (argmax)

ry · ID 541 · rank / rang 1

Logit 33.729183197021484 · P 0.9999608664391816 · log(P) -0.00003913432655622273

The most likely for this prefix. It does not decide which token is appended in our example.

Imposed choice — not selected by the model

ry · ID 541 · rank / rang 1

Logit 33.729183197021484 · P 0.9999608664391816 · log(P) -0.00003913432655622273

Target already in the top 8; its probability is counted only once.

Imposed choice — not selected by the model · ry · ID 541
From representation to probabilities

d0: -2.991468667984009 d1: -0.8253179788589478 d2: 2.7496345043182373 d3: 1.4738621711730957 d4: -2.201028823852539 d5: -0.8349835276603699 d6: 1.1787775754928589 d7: -0.9351224899291992

Top 8 candidates and imposed target. Fixture values displayed by the browser, scientific notation when needed; no renormalization. Full-vocabulary logsumexp: 33.72922233134804

  • ry · ID 541 · rank / rang 1

    Logit 33.729183197021484 · P 0.9999608664391816 · log(P) -0.00003913432655622273

  • r · ID 98 · rank / rang 2

    Logit 23.45769500732422 · P 0.000034604482785306724 · log(P) -10.271527324023822

  • ries · ID 1333 · rank / rang 3

    Logit 20.316591262817383 · P 0.0000014961265817374285 · log(P) -13.412631068530658

  • os · ID 395 · rank / rang 4

    Logit 19.061750411987305 · P 4.265774383831879e-7 · log(P) -14.667471919360736

  • rell · ID 12159 · rank / rang 5

    Logit 18.704822540283203 · P 2.985286911238753e-7 · log(P) -15.024399791064836

  • rent · ID 1160 · rank / rang 6

    Logit 18.517757415771484 · P 2.475966278494743e-7 · log(P) -15.211464915576556

  • rying · ID 11089 · rank / rang 7

    Logit 18.411640167236328 · P 2.226684007938849e-7 · log(P) -15.317582164111712

  • ried · ID 2041 · rank / rang 8

    Logit 18.383697509765625 · P 2.1653257875596153e-7 · log(P) -15.345524821582416

Other tokens : 0.0000016210477123238883 · 49144 omitted candidates.

These tokens are neither in the top 8 nor the imposed target. The target’s probability counts once, even when it appears in both panels. other_probability = mass outside UNION(top8, target).

The maximum breaks ties by lowest ID. The script imposes the target, even when they coincide. Temperature 1, no filters or sampling.

The ordinary loop, unlike our example

In ordinary autoregressive generation, a decoding rule chooses the next token from scores for the current context, extends that context, and repeats. It does not require a complete reply written beforehand.

Hi, I will be a little late. Sorry
g=11 → 12 · ID 288
<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
<|im_start|>user
Write a message to let someone know I will be late.<|im_end|>
<|im_start|>assistant
Hi, I will be a little late. Sorry

1, 9690, 198, 2683, 359, 253, 5356, 5646, 11173, 3365, 3511, 308, 34519, 28, 7018, 411, 407, 19712, 8182, 2, 198, 1, 4093, 198, 19161, 253, 3714, 288, 1303, 2206, 699, 339, 523, 325, 3008, 30, 2, 198, 1, 520, 9531, 198, 26843, 28, 339, 523, 325, 253, 1838, 3008, 30, 25442, 541

Recorded model calculations — given the displayed prefix · Step prefix 11 · Documented reduced view

Final representation at the last position → logits → full-vocabulary softmax.

Probability ≠ truth. These scores concern continuation of the prefix, not the reliability of a claim.

Model maximum (argmax)

Ġfor · ID 327 · rank / rang 1

Logit 17.719886779785156 · P 0.4803608658735202 · log(P) -0.7332176536400254

The most likely for this prefix. It does not decide which token is appended in our example.

Imposed choice — not selected by the model

Ġto · ID 288 · rank / rang 4

Logit 15.189250946044922 · P 0.038240753466153105 · log(P) -3.26385348738026

Target already in the top 8; its probability is counted only once.

Imposed choice — not selected by the model · Ġto · ID 288
From representation to probabilities

d0: 0.29135775566101074 d1: -2.2683322429656982 d2: 0.8311927318572998 d3: 1.8787803649902344 d4: 1.275809407234192 d5: -2.0554420948028564 d6: -1.340942621231079 d7: -1.1240665912628174

Top 8 candidates and imposed target. Fixture values displayed by the browser, scientific notation when needed; no renormalization. Full-vocabulary logsumexp: 18.45310443342518

  • Ġfor · ID 327 · rank / rang 1

    Logit 17.719886779785156 · P 0.4803608658735202 · log(P) -0.7332176536400254

  • Ġabout · ID 563 · rank / rang 2

    Logit 17.004520416259766 · P 0.2349026708638545 · log(P) -1.448584017165416

  • , · ID 28 · rank / rang 3

    Logit 16.46103286743164 · P 0.13641254495466326 · log(P) -1.992071565993541

  • Ġto · ID 288 · rank / rang 4

    Logit 15.189250946044922 · P 0.038240753466153105 · log(P) -3.26385348738026

  • ! · ID 17 · rank / rang 5

    Logit 14.969992637634276 · P 0.03071169366446069 · log(P) -3.4831117957909044

  • . · ID 30 · rank / rang 6

    Logit 14.595712661743164 · P 0.02112302139961681 · log(P) -3.8573917716820176

  • ĠI · ID 339 · rank / rang 7

    Logit 14.11510944366455 · P 0.013062692797345517 · log(P) -4.337994989760631

  • Ġif · ID 585 · rank / rang 8

    Logit 13.74203109741211 · P 0.008995117615490225 · log(P) -4.711073336013072

Other tokens : 0.03619063936489409 · 49144 omitted candidates.

These tokens are neither in the top 8 nor the imposed target. The target’s probability counts once, even when it appears in both panels. other_probability = mass outside UNION(top8, target).

The maximum breaks ties by lowest ID. The script imposes the target, even when they coincide. Temperature 1, no filters or sampling.

The ordinary loop, unlike our example

In ordinary autoregressive generation, a decoding rule chooses the next token from scores for the current context, extends that context, and repeats. It does not require a complete reply written beforehand.

Hi, I will be a little late. Sorry to
g=12 → 13 · ID 1446
<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
<|im_start|>user
Write a message to let someone know I will be late.<|im_end|>
<|im_start|>assistant
Hi, I will be a little late. Sorry to

1, 9690, 198, 2683, 359, 253, 5356, 5646, 11173, 3365, 3511, 308, 34519, 28, 7018, 411, 407, 19712, 8182, 2, 198, 1, 4093, 198, 19161, 253, 3714, 288, 1303, 2206, 699, 339, 523, 325, 3008, 30, 2, 198, 1, 520, 9531, 198, 26843, 28, 339, 523, 325, 253, 1838, 3008, 30, 25442, 541, 288

Recorded model calculations — given the displayed prefix · Step prefix 12 · Documented reduced view

Final representation at the last position → logits → full-vocabulary softmax.

Probability ≠ truth. These scores concern continuation of the prefix, not the reliability of a claim.

Model maximum (argmax)

Ġhear · ID 4875 · rank / rang 1

Logit 19.18226623535156 · P 0.3051657244781808 · log(P) -1.18690029099589

The most likely for this prefix. It does not decide which token is appended in our example.

Imposed choice — not selected by the model

Ġkeep · ID 1446 · rank / rang 21

Logit 15.25147533416748 · P 0.00598983632353836 · log(P) -5.117691192179972

Target outside top 8 — not promoted in the ranking.

Imposed choice — not selected by the model · Ġkeep · ID 1446
From representation to probabilities

d0: -0.5657979249954224 d1: -1.826673150062561 d2: -1.1230062246322632 d3: 0.407820463180542 d4: 0.2504291236400604 d5: -0.6481521129608154 d6: -1.0134414434432983 d7: -2.165100574493408

Top 8 candidates and imposed target. Fixture values displayed by the browser, scientific notation when needed; no renormalization. Full-vocabulary logsumexp: 20.369166526347453

  • Ġhear · ID 4875 · rank / rang 1

    Logit 19.18226623535156 · P 0.3051657244781808 · log(P) -1.18690029099589

  • Ġbreak · ID 2373 · rank / rang 2

    Logit 18.23504638671875 · P 0.11834867488903218 · log(P) -2.1341201396287026

  • Ġbother · ID 18531 · rank / rang 3

    Logit 17.561723709106445 · P 0.060359144574573455 · log(P) -2.8074428172410073

  • Ġbump · ID 22908 · rank / rang 4

    Logit 17.550188064575195 · P 0.05966686356921321 · log(P) -2.8189784617722573

  • Ġthink · ID 1510 · rank / rang 5

    Logit 17.540145874023438 · P 0.05907067608004241 · log(P) -2.829020652324015

  • Ġget · ID 820 · rank / rang 6

    Logit 16.91200065612793 · P 0.031518964388736165 · log(P) -3.457165870219523

  • Ġdisturb · ID 7297 · rank / rang 7

    Logit 16.869901657104492 · P 0.030219590609247338 · log(P) -3.4992648692429604

  • Ġbe · ID 325 · rank / rang 8

    Logit 16.679821014404297 · P 0.02498835126897151 · log(P) -3.689345511943156

Other tokens : 0.3046721738184629 · 49143 omitted candidates.

These tokens are neither in the top 8 nor the imposed target. The target’s probability counts once, even when it appears in both panels. other_probability = mass outside UNION(top8, target).

The maximum breaks ties by lowest ID. The script imposes the target, even when they coincide. Temperature 1, no filters or sampling.

The ordinary loop, unlike our example

In ordinary autoregressive generation, a decoding rule chooses the next token from scores for the current context, extends that context, and repeats. It does not require a complete reply written beforehand.

Hi, I will be a little late. Sorry to keep
g=13 → 14 · ID 346
<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
<|im_start|>user
Write a message to let someone know I will be late.<|im_end|>
<|im_start|>assistant
Hi, I will be a little late. Sorry to keep

1, 9690, 198, 2683, 359, 253, 5356, 5646, 11173, 3365, 3511, 308, 34519, 28, 7018, 411, 407, 19712, 8182, 2, 198, 1, 4093, 198, 19161, 253, 3714, 288, 1303, 2206, 699, 339, 523, 325, 3008, 30, 2, 198, 1, 520, 9531, 198, 26843, 28, 339, 523, 325, 253, 1838, 3008, 30, 25442, 541, 288, 1446

Recorded model calculations — given the displayed prefix · Step prefix 13 · Documented reduced view

Final representation at the last position → logits → full-vocabulary softmax.

Probability ≠ truth. These scores concern continuation of the prefix, not the reliability of a claim.

Model maximum (argmax)

Ġyou · ID 346 · rank / rang 1

Logit 19.786155700683594 · P 0.9542618840985152 · log(P) -0.04681713357161144

The most likely for this prefix. It does not decide which token is appended in our example.

Imposed choice — not selected by the model

Ġyou · ID 346 · rank / rang 1

Logit 19.786155700683594 · P 0.9542618840985152 · log(P) -0.04681713357161144

Target already in the top 8; its probability is counted only once.

Imposed choice — not selected by the model · Ġyou · ID 346
From representation to probabilities

d0: -0.5804969668388367 d1: -2.092398166656494 d2: -0.4034739136695862 d3: -0.2988728880882263 d4: -1.1881353855133057 d5: -0.13915109634399414 d6: 0.38976895809173584 d7: -0.5138341784477234

Top 8 candidates and imposed target. Fixture values displayed by the browser, scientific notation when needed; no renormalization. Full-vocabulary logsumexp: 19.832972834255205

  • Ġyou · ID 346 · rank / rang 1

    Logit 19.786155700683594 · P 0.9542618840985152 · log(P) -0.04681713357161144

  • Ġeveryone · ID 2573 · rank / rang 2

    Logit 14.890088081359863 · P 0.00713398886490547 · log(P) -4.942884752895342

  • Ġon · ID 335 · rank / rang 3

    Logit 14.530855178833008 · P 0.004981034615424339 · log(P) -5.302117655422197

  • Ġpeople · ID 701 · rank / rang 4

    Logit 14.043898582458496 · P 0.0030608144086935237 · log(P) -5.789074251796709

  • Ġit · ID 357 · rank / rang 5

    Logit 13.720077514648438 · P 0.0022141309068167787 · log(P) -6.112895319606768

  • Ġme · ID 549 · rank / rang 6

    Logit 13.581164360046388 · P 0.001926966114660211 · log(P) -6.2518084742088185

  • Ġto · ID 288 · rank / rang 7

    Logit 13.44144344329834 · P 0.0016756914532827397 · log(P) -6.391529390956865

  • Ġu · ID 326 · rank / rang 8

    Logit 13.291946411132812 · P 0.0014430066032812478 · log(P) -6.541026423122393

Other tokens : 0.023302482934420325 · 49144 omitted candidates.

These tokens are neither in the top 8 nor the imposed target. The target’s probability counts once, even when it appears in both panels. other_probability = mass outside UNION(top8, target).

The maximum breaks ties by lowest ID. The script imposes the target, even when they coincide. Temperature 1, no filters or sampling.

The ordinary loop, unlike our example

In ordinary autoregressive generation, a decoding rule chooses the next token from scores for the current context, extends that context, and repeats. It does not require a complete reply written beforehand.

Hi, I will be a little late. Sorry to keep you
g=14 → 15 · ID 7003
<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
<|im_start|>user
Write a message to let someone know I will be late.<|im_end|>
<|im_start|>assistant
Hi, I will be a little late. Sorry to keep you

1, 9690, 198, 2683, 359, 253, 5356, 5646, 11173, 3365, 3511, 308, 34519, 28, 7018, 411, 407, 19712, 8182, 2, 198, 1, 4093, 198, 19161, 253, 3714, 288, 1303, 2206, 699, 339, 523, 325, 3008, 30, 2, 198, 1, 520, 9531, 198, 26843, 28, 339, 523, 325, 253, 1838, 3008, 30, 25442, 541, 288, 1446, 346

Recorded model calculations — given the displayed prefix · Step prefix 14 · Documented reduced view

Final representation at the last position → logits → full-vocabulary softmax.

Probability ≠ truth. These scores concern continuation of the prefix, not the reliability of a claim.

Model maximum (argmax)

Ġwaiting · ID 7003 · rank / rang 1

Logit 26.402332305908203 · P 0.858207166551665 · log(P) -0.1529097557818062

The most likely for this prefix. It does not decide which token is appended in our example.

Imposed choice — not selected by the model

Ġwaiting · ID 7003 · rank / rang 1

Logit 26.402332305908203 · P 0.858207166551665 · log(P) -0.1529097557818062

Target already in the top 8; its probability is counted only once.

Imposed choice — not selected by the model · Ġwaiting · ID 7003
From representation to probabilities

d0: 0.2854216992855072 d1: -2.1910338401794434 d2: -0.1800287663936615 d3: 1.2428923845291138 d4: 0.29130080342292786 d5: -1.5335160493850708 d6: 0.180755153298378 d7: -0.8591192960739136

Top 8 candidates and imposed target. Fixture values displayed by the browser, scientific notation when needed; no renormalization. Full-vocabulary logsumexp: 26.55524206169001

  • Ġwaiting · ID 7003 · rank / rang 1

    Logit 26.402332305908203 · P 0.858207166551665 · log(P) -0.1529097557818062

  • Ġup · ID 614 · rank / rang 2

    Logit 24.415771484375 · P 0.1177171486562268 · log(P) -2.1394705773150093

  • Ġon · ID 335 · rank / rang 3

    Logit 22.057783126831055 · P 0.011137261117958024 · log(P) -4.497458934858955

  • Ġso · ID 588 · rank / rang 4

    Logit 20.341081619262695 · P 0.0020008955124155513 · log(P) -6.214160442427314

  • Ġin · ID 281 · rank / rang 5

    Logit 20.16312789916992 · P 0.0016747118430241522 · log(P) -6.3921141625200875

  • Ġawaiting · ID 28121 · rank / rang 6

    Logit 19.40749740600586 · P 0.0007866362176931138 · log(P) -7.14774465568415

  • Ġoff · ID 767 · rank / rang 7

    Logit 19.16392517089844 · P 0.0006165834455731841 · log(P) -7.391316890791572

  • Ġawake · ID 22576 · rank / rang 8

    Logit 19.057186126708984 · P 0.0005541606479624322 · log(P) -7.498055934981025

Other tokens : 0.007305436007481138 · 49144 omitted candidates.

These tokens are neither in the top 8 nor the imposed target. The target’s probability counts once, even when it appears in both panels. other_probability = mass outside UNION(top8, target).

The maximum breaks ties by lowest ID. The script imposes the target, even when they coincide. Temperature 1, no filters or sampling.

The ordinary loop, unlike our example

In ordinary autoregressive generation, a decoding rule chooses the next token from scores for the current context, extends that context, and repeats. It does not require a complete reply written beforehand.

Hi, I will be a little late. Sorry to keep you waiting
g=15 → 16 · ID 17
<|im_start|>system
You are a helpful AI assistant named SmolLM, trained by Hugging Face<|im_end|>
<|im_start|>user
Write a message to let someone know I will be late.<|im_end|>
<|im_start|>assistant
Hi, I will be a little late. Sorry to keep you waiting

1, 9690, 198, 2683, 359, 253, 5356, 5646, 11173, 3365, 3511, 308, 34519, 28, 7018, 411, 407, 19712, 8182, 2, 198, 1, 4093, 198, 19161, 253, 3714, 288, 1303, 2206, 699, 339, 523, 325, 3008, 30, 2, 198, 1, 520, 9531, 198, 26843, 28, 339, 523, 325, 253, 1838, 3008, 30, 25442, 541, 288, 1446, 346, 7003

Recorded model calculations — given the displayed prefix · Step prefix 15 · Documented reduced view

Final representation at the last position → logits → full-vocabulary softmax.

Probability ≠ truth. These scores concern continuation of the prefix, not the reliability of a claim.

Model maximum (argmax)

. · ID 30 · rank / rang 1

Logit 22.8670654296875 · P 0.6672956779003861 · log(P) -0.4045220360898263

The most likely for this prefix. It does not decide which token is appended in our example.

Imposed choice — not selected by the model

! · ID 17 · rank / rang 3

Logit 20.82699203491211 · P 0.0867612287175501 · log(P) -2.444595430865217

Target already in the top 8; its probability is counted only once.

Imposed choice — not selected by the model · ! · ID 17
From representation to probabilities

d0: 0.33285361528396606 d1: -1.1384018659591677 d2: -0.47933995723724365 d3: 2.0314106941223145 d4: 1.1137346029281616 d5: -0.7825533747673035 d6: -2.0709948539733887 d7: -1.766234278678894

Top 8 candidates and imposed target. Fixture values displayed by the browser, scientific notation when needed; no renormalization. Full-vocabulary logsumexp: 23.271587465777326

  • . · ID 30 · rank / rang 1

    Logit 22.8670654296875 · P 0.6672956779003861 · log(P) -0.4045220360898263

  • , · ID 28 · rank / rang 2

    Logit 21.341018676757812 · P 0.14506566325691453 · log(P) -1.9305687890195136

  • ! · ID 17 · rank / rang 3

    Logit 20.82699203491211 · P 0.0867612287175501 · log(P) -2.444595430865217

  • Ġfor · ID 327 · rank / rang 4

    Logit 20.000686645507812 · P 0.03797220553133754 · log(P) -3.270900820269514

  • Ġso · ID 588 · rank / rang 5

    Logit 19.153820037841797 · P 0.016280822005639187 · log(P) -4.117767427935529

  • Ġand · ID 284 · rank / rang 6

    Logit 18.485294342041016 · P 0.008343327813971972 · log(P) -4.786293123736311

  • Ġbut · ID 564 · rank / rang 7

    Logit 17.500877380371094 · P 0.0031175430062211373 · log(P) -5.7707100854062325

  • Ġa · ID 253 · rank / rang 8

    Logit 17.101383209228516 · P 0.0020908088954878805 · log(P) -6.170204256548811

Other tokens : 0.03307272287249243 · 49144 omitted candidates.

These tokens are neither in the top 8 nor the imposed target. The target’s probability counts once, even when it appears in both panels. other_probability = mass outside UNION(top8, target).

The maximum breaks ties by lowest ID. The script imposes the target, even when they coincide. Temperature 1, no filters or sampling.

The ordinary loop, unlike our example

In ordinary autoregressive generation, a decoding rule chooses the next token from scores for the current context, extends that context, and repeats. It does not require a complete reply written beforehand.

Hi, I will be a little late. Sorry to keep you waiting!

No KV cache in this capture; the visit replays prefixes, it does not recalculate the model.

Recorded calculations for successive prefixes; tokens imposed by the script. End of prepared path — not a model-selected stop.

Go deeper

In ordinary autoregressive generation, a decoding rule chooses the next token from scores for the current context, extends that context, and repeats. It does not require a complete reply written beforehand.

The scored prefix contains the initial input and only the tokens already imposed. No KV cache in this capture. The prepared ending is neither a chosen EOS nor a reached length limit.

Cette scène en français

S09 · The prepared reply returns to the screen

Prepared message : Write a message to let someone know I will be late.

Is each bit appearing on screen one new model token?

Prepared reply — decoded from the imposed path. The model did not freely choose this text; the return is staged.

Generating tokens, transporting data and showing text are different steps. Their timing is simplified in this demonstration.

  1. You, at the same deskThe journey returns to its starting point.
  2. ApplicationDisplays the decoded text.
  3. Prepared replyWritten by our team, not freely chosen by the model.
The return displays; it no longer generates. Character ≠ token ≠ network packet.◇ Teaching diagram · illustrative shapes and distances, not a model measurement.

Prepared reply — decoded from the imposed path. The model did not freely choose this text; the return is staged.

Show me

Must a network chunk, a token and a visible character be the same thing?

Prepared reply — decoded from the imposed path. The model did not freely choose this text; the return is staged.

Prepared reply — written by our team

Hi, I will be a little late. Sorry to keep you waiting!

End of prepared path — not a model-selected stop: prepared_sequence_end · 16 tokens. No EOS appended, no invented terminal distribution.

Output IDs and decoding

26843, 28, 339, 523, 325, 253, 1838, 3008, 30, 25442, 541, 288, 1446, 346, 7003, 17

Hi, I will be a little late. Sorry to keep you waiting!

Decoding: skip_special_tokens=false; clean_up_tokenization_spaces=false. No special token appended or measured network chunks. Character ≠ token ≠ packet.

Formatting and transport need not deliver one token at a time. A visible character may require more than one token.

Go deeper

The exact decoder options, handling of special tokens and any normalization belong to the capture manifest. A readable view must retain access to the raw IDs and decoded text, including spaces and line breaks.

The return-to-screen animation is not a network recording or a latency measurement. It may replay an already completed output; it must not claim the model is still computing.

Cette scène en français

S10 · Explain the mechanism and our exception

Prepared message : Write a message to let someone know I will be late.

What would ordinary generation do differently from our authored example?

The application prepares an input. A trained model calculates possible next tokens; a decoding rule chooses one, adds it to the context, and repeats. The application displays the resulting text.

The imposed-path demonstration does not establish autonomous generation or factual truth.

  1. The application preparesMessage, envelope and tokens.
  2. The model calculatesRepresentations, then conditional scores.
  3. The application displaysHere, the tokens of the reply written by our team.
An explainable mechanism, with limits. The five distinctions below hold even when the sentence sounds plausible.◇ Teaching diagram · illustrative shapes and distances, not a model measurement.

The imposed-path demonstration does not establish autonomous generation or factual truth.

Show me

What has to happen before the next token can be chosen?

Five distinctions to keep

  • Server ≠ model

    The computer hosts the software that runs the learned calculation.

    Revisit S02
  • Tokenization ≠ understanding

    Splitting text and assigning IDs prepares calculation; it does not prove understanding.

    Revisit S04
  • Training ≠ generation

    Learning adjusts parameters before the conversation. They stay fixed during this reply.

    Revisit S03
  • Probability ≠ truth

    A high continuation score does not check facts. The maximum is not a guarantee.

    Revisit S07
  • Guided ≠ autonomous

    Our team wrote the reply and imposed its tokens. Recorded conditional calculations do not prove the model would have chosen this text.

    Revisit S08

The journey does not complete any lesson for you. Continue with the six existing lessons. Open the lessons

Explain it aloud or to yourself. We do not collect or save your explanation.

Before the model processes the request…

The input is prepared now; the learned parameters already exist. See S02–S03.

Before selecting the next token…

Scores concern continuations under this context, not factual verification. See S06–S07.

In ordinary generation, after a token is chosen…

The token extends the prefix for the next calculation. Only in our example does the script impose that token. See S08.

When the reply appears…

Display is not a truth test, and its timing is not the generation timing. See S09–S10.

Show a reference explanation

In ordinary autoregressive generation, a decoding rule chooses the next token from scores for the current context, extends that context, and repeats. It does not require a complete reply written beforehand. Recorded calculations for successive prefixes; tokens imposed by the script. End of prepared path — not a model-selected stop. The imposed-path demonstration does not establish autonomous generation or factual truth.

The imposed-path demonstration does not establish autonomous generation or factual truth.

Go deeper

A model is not the entire assistant. An application may add search, tools or external memory, with their own behaviour and permissions. Only show such capabilities as used when the trace actually contains them.

The six existing lessons and their labs remain available now. They deepen generation, tokens, variation, hallucinations, context and systems; their prepared simulations are not this model's captured trace.

Cette scène en français

This is the complete reading version. Interactive controls require JavaScript; the explanations and sources do not.

Existing site · Explore the six lessons

Explore the six lessons

Provenance, sources and limits of this capture

Introduction of the original editorial proposal, now bound to verified data:

Our team wrote the example reply to keep the story easy to follow. This proposed journey would show model calculations along that imposed path, not a freely generated reply. No model is called during a visit.

In ordinary autoregressive generation, a decoding rule chooses the next token from scores for the current context, extends that context, and repeats. It does not require a complete reply written beforehand.

HuggingFaceTB/SmolLM2-135M-Instruct · 12fd25f77366fa6b3b4b768ec3050bf629380bac · Apache-2.0 · float32 · CPU · PyTorch 2.8.0+cpu · Transformers 4.57.6. Reproducibility seed 0, no sampling. Not a live capture.

SHA256 5af571cbf074e6d21a03528d2330792e532ca608f24ac70a143f6b369968ab8c

journey-guided-trace-v1 · late-authored-v1 · en

Trace manifest · Model / tokenizer

One head and one layer of the initial prompt, eight selected dimensions; no weights shipped. Diagrams and camera do not represent measured transport or latency.