View on GitHub

Tokenization: why AI can’t count letters

AI engineering planned

Ask an AI how many r’s are in ‘strawberry’. Here’s why it struggles.

The idea

Language models don’t read letters. They read tokens: chunks of text from a fixed vocabulary, where a common word might be one token and a rare word several. ‘strawberry’ may arrive as a few pieces, so the model never sees the individual letters it’s asked to count.

Tokens also explain why models price by the token, why some languages cost more, and why numbers behave oddly. This lesson builds byte-pair encoding, the algorithm behind most tokenizers, from scratch.

What the lesson will build

Key ideas

The video

When it’s published, the code will live in ai-engineering/ and this page will link to it.


All topics · Suggest a topic