Current generation
Roughly 200,000 merges, and the encoding behind every current OpenAI model — the GPT-5 family, GPT-4.1 and GPT-4o included. Its larger vocabulary compresses non-English text far better than its predecessor: Hindi, Chinese and Japanese cost roughly half the tokens they did under cl100k_base. The newest registry entry, o200k_harmony, reuses these exact ranks and only adds control tokens, so plain-text counts are identical.