Computers can't read letters — only numbers. So the first thing any language model does is chop your text into pieces called tokens: whole words, word chunks, sometimes single letters. Each piece gets an ID number. From this point on, the model never sees your words — only the numbers.
Why chop text at all? Because the model only speaks numbers — and the vocabulary size is a tradeoff. Too few tiles and every sentence becomes a long crawl of letters; too many and the final layer, which scores every tile, gets enormous and slow. Real vocabularies land around 100,000–200,000 tiles: the compromise between short sequences and a manageable output layer.
The standard recipe for learning those tiles is BPE — byte-pair encoding: start from single characters, repeatedly merge the most frequent neighboring pair, stop at the target size. The playground below runs the genuine algorithm, trained live on a toy corpus in your browser.
What do the little numbers mean — like the 41 stamped on “un”? Each one is the tile's address in the vocabulary: one giant numbered list of every tile the model knows. “un” is entry #41, so your text reaches the model as the number 41 — it never sees the letters u-n at all. The numbers themselves are arbitrary (another model's “un” might be #9,204); what matters is that words become a list of numbers the model can compute with. In the next stage, each number gets traded in for its starting meaning.
Simplified illustration · live toy model
Tokenizer playground
A genuine BPE tokenizer — “byte-pair encoding”, the standard trick tokenizers use to learn their pieces — trained live in your browser on 49 tiny sentences. Toy vocabulary: 75 pieces. GPT's real one learns exactly this way, at vastly larger scale.
Merge trace for “unbelievable” — every BPE merge, in order
Fifty tokenizations, up close
How close is the toy to the real thing? Closer than it looks: the playground above runs the genuine BPE merge loop — count adjacent pairs, merge the most frequent, repeat — the same core algorithm inside GPT-2's tokenizer, tiktoken's BPE mode, and SentencePiece's BPE mode. The tokenizer family differs in details, not in the idea. Six of them, side by side:
The tokenizer family
Why should you care how text gets chopped? Three ways it bites in real life. The model can’t truly see spelling — “the” is one tile but “teh” is three, so a typo looks like a different word entirely. It can’t do arithmetic — numbers arrive as digit confetti, never as values. And other languages pay more — English gets roughly a tile per word while Chinese often costs a tile per character, making the same meaning longer, slower, and pricier to process.
Most tokenizers learn their vocabulary with byte-pair encoding: start from single characters and repeatedly merge the most frequent adjacent pair, until the vocabulary hits its target size — tens of thousands of pieces. Frequent words earn their own token; rare words get spelled out in chunks, so any text can be represented. Rule of thumb: about 4 characters ≈ 1 token in English. The tokenizer is frozen before training and shared by the model forever — which is why a model can never learn a new word's spelling after the fact.