Tokenizers: how text becomes token IDs
A tokenizer is the component that converts text into the numbered tokens a language model works with, and converts the model's tokens back into text.
What a tokenizer is
Before a model sees your prompt, a tokenizer chops it into tokens from a fixed vocabulary and replaces each with an integer ID. When the model replies, the same tokenizer turns its IDs back into text. Each model family ships with its own tokenizer, trained alongside the model, and the two cannot be swapped.
How it works
Most modern tokenizers use byte pair encoding (BPE) or a close relative such as the SentencePiece unigram method. BPE starts from bytes or characters and repeatedly merges the most frequent adjacent pairs in a training corpus until it reaches the target vocabulary size, often 100,000 to 200,000 entries in current models. The result:
- Frequent words become single tokens, usually with their leading space attached.
- Rare words split into several subword pieces.
- "Hello", " Hello" and " hello" are typically three different tokens.
- Special tokens mark structure, such as the start and end of a turn. These are what an instruct template is built from.
Byte-level schemes mean any input can be encoded, including emoji and unusual scripts, but uncommon text costs more tokens.
Why it matters in roleplay setups
- Token counting. SillyTavern estimates prompt size with a selected tokenizer. If it is the wrong one, it may overfill the context and the API rejects the request, or underfill it and waste history.
- Logit bias. Bans are keyed by token ID, so they only work with the model's own tokenizer.
- Names and invented words. A fantasy name split into four tokens is slightly harder for the model to reproduce exactly, and token-level penalties can break it into odd spellings.
- Cost. The same text costs a different number of tokens on different model families.
Common mistakes
Leaving the frontend tokenizer on a default that belongs to another model family; treating a token count from one model as exact for another; and building logit bias lists by hand from a token ID table for a different tokenizer. When in doubt, the usage numbers returned by the API are the real count.
Tokenizers on Wild West API
Wild West API counts and bills tokens with each model's own tokenizer, and the counts come back in the response usage field. With chat completions, the server also applies the model's chat template, so you never need to insert special tokens yourself. In SillyTavern, choosing a tokenizer that matches the model family, or the best match available, keeps its estimates close.
FAQ
Do all LLMs use the same tokenizer?
No. Each model family has its own vocabulary, so the same text produces different token counts and IDs.
Which tokenizer should I pick in SillyTavern?
The one matching your model's family if it is listed. Otherwise pick a close one and check the API's usage numbers to see how far off it is.