An interactive blueprint · tap any stage

How a large
language model works

A language model does one trick: it guesses the next word. To pull it off, it turns your words into numbers, runs them through a stack of identical blocks that mix in context, then rolls weighted dice to pick what comes next — one word at a time, hundreds of times per answer. Tap any stage below and play with the working toys.

raw text tokenizer token IDs embeddings transformerblocks ×N logits sampling text out append token & repeat
One token’s journey: text → numbers → the machine → a dice roll → text, then the loop repeats.

Separate track

How it learned

Everything above is the performance. This is the rehearsal. Nobody programmed any of it — every number in the machine started as random noise, and training shaped them in four phases. Same prompt, four different behaviors:

web · books · code trillions of tokens A · pretraining predict next token → base model instruction demos Q → good answer B · fine-tuning learn from demos (SFT) → assistant math · code · tools answers that check C · reasoning RL reward what checks out → thinking model human rankings best → worst D · RLHF prefer what humans like → aligned assistant
Training in four phases: raw capability, then shape, then reasoning, then taste.

One prompt, four trainings

Simplified illustration · same prompt: “Why is the sky blue?”

“Why is the sky blue? Great question! Here are 10 more questions about the sky: 1. Why is the ocean blue? 2. What makes sunsets red? 3. …”

After pretraining: a brilliant mimic, not an assistant. It continues the text instead of answering — it never learned that questions want answers.

1 / 4 · pretraining

Simplified illustration · same prompt: “Why is the sky blue?”

“Sunlight scatters off air molecules, and blue light scatters the most — so blue arrives at your eyes from every direction. That's the short version!”

After fine-tuning: now it knows the shape of helpful — answer the question, keep it tight, hold a friendly tone.

2 / 4 · fine-tuning

Simplified illustration · same prompt: “Why is the sky blue?”

“Let me think it through: sunlight hits air molecules → Rayleigh scattering, strength ∝ 1/λ⁴ → blue (short λ) scatters ~5× more than red → blue comes from all directions. Checking: that also explains red sunsets (blue scattered away). Confident.”

After reasoning RL: it learned to show its work — because attempts with checkable reasoning got rewarded. The thinking is the product.

3 / 4 · reasoning RL

Simplified illustration · same prompt: “Why is the sky blue?”

“Great question! In short: air scatters blue sunlight more than other colors, so the whole sky glows blue. (That's the simplified version — happy to go deeper into the physics if you'd like!)”

After RLHF: same facts, better manners — warm, humble about simplifications, invites follow-up. Tuned toward what humans prefer.

4 / 4 · RLHF