Computers can't read letters — only numbers. So the first thing any language model does is chop your text into pieces called tokens: whole words, word chunks, sometimes single letters. Each piece gets an ID number. From this point on, the model never sees your words — only the numbers.
Why chop text at all? Because the model only speaks numbers — and the vocabulary size is a tradeoff. Too few tiles and every sentence becomes a long crawl of letters; too many and the final layer, which scores every tile, gets enormous and slow. Real vocabularies land around 100,000–200,000 tiles: the compromise between short sequences and a manageable output layer.
The standard recipe for learning those tiles is BPE — byte-pair encoding: start from single characters, repeatedly merge the most frequent neighboring pair, stop at the target size. The playground below runs the genuine algorithm, trained live on a toy corpus in your browser.
Simplified illustration · live toy model
Tokenizer playground
A genuine BPE tokenizer — “byte-pair encoding”, the standard trick tokenizers use to learn their pieces — trained live in your browser on 49 tiny sentences. Toy vocabulary: 75 pieces. GPT's real one learns exactly this way, at vastly larger scale.
Merge trace for “unbelievable” — every BPE merge, in order
Fifty tokenizations, up close
Why should you care how text gets chopped? Three ways it bites in real life. The model can’t truly see spelling — “the” is one tile but “teh” is three, so a typo looks like a different word entirely. It can’t do arithmetic — numbers arrive as digit confetti, never as values. And other languages pay more — English gets roughly a tile per word while Chinese often costs a tile per character, making the same meaning longer, slower, and pricier to process.
Most tokenizers learn their vocabulary with byte-pair encoding: start from single characters and repeatedly merge the most frequent adjacent pair, until the vocabulary hits its target size — tens of thousands of pieces. Frequent words earn their own token; rare words get spelled out in chunks, so any text can be represented. Rule of thumb: about 4 characters ≈ 1 token in English. The tokenizer is frozen before training and shared by the model forever — which is why a model can never learn a new word's spelling after the fact.