Different models read text with different tokenizers, and the differences are bigger than you’d think: emojis, code, numbers, and non-English text can cost very different amounts on different models. Paste anything and compare, piece by piece.

The left panel is o200k_base (GPT-4o, GPT-4o mini, o1); the right is cl100k_base (GPT-4, GPT-3.5, and the OpenAI embedding models). Fewer tokens for the same text means lower cost and more room in the context window, which is why newer tokenizers are usually leaner.

Want the price attached? The token & cost calculator turns these counts into dollars across models.