freecoding.school100% FREE · NO SIGNUP
Tensor TownISSUE #51 of 120

tokenization · bpe · wordpiece · unigram

NeuraVSThe Overfit Ogre
Neura saysTokenization splits text into tokens — subword pieces the model actually processes.

Models don't see characters or words; they see tokens. A tokenizer (BPE, WordPiece, Unigram) splits text into common subword chunks — frequent words are one token, rare words split into pieces. This balances vocabulary size against sequence length. Token counts drive cost and context limits, and quirks of tokenization explain odd model behaviors (e.g. counting letters, arithmetic on digits).

Power-ups you unlock

The Overfit Ogre attacks — common mistakes

Boss battleExplain why a model may struggle to count the letters in a word.

Example code

<!doctype html><html><head><meta charset="utf-8"></head>
<body style="background:#06040d;color:#e6e0ff;font-family:monospace;padding:20px"><pre>"tokenization" → ["token","ization"] (2 tokens)
model sees tokens, not letters
→ letter-counting is hard</pre></body></html>
▶ Open the interactive comic issue
‹ Encoder-Decoder · T5 · BartContext Windows · How Much Fits At Once ›