Models don't see characters or words; they see tokens. A tokenizer (BPE, WordPiece, Unigram) splits text into common subword chunks — frequent words are one token, rare words split into pieces. This balances vocabulary size against sequence length. Token counts drive cost and context limits, and quirks of tokenization explain odd model behaviors (e.g. counting letters, arithmetic on digits).