Wild West API

Tokens: the units LLMs read, write and bill by

A token is a small chunk of text, often a word or part of a word, that a language model reads and writes one at a time.

What a token is

Models do not see letters or words directly. A tokenizer splits text into pieces from a fixed vocabulary, and each piece has an ID number. Common words are usually one token, often including the space before them (" the"). Rare words, names and made-up fantasy terms get split into several pieces. Punctuation, emoji and formatting characters are tokens too.

As a rough rule for English, one token is about four characters, or about three quarters of a word, so 1,000 tokens is roughly 750 words. Other languages, code and heavy formatting often use more tokens per word.

Why tokens matter

  • Billing. APIs charge per token, with input (everything you send) and output (what the model writes) usually priced separately.
  • Limits. The context window and max tokens are measured in tokens, not words or characters.
  • Sampling. Samplers like temperature and min P act on each token choice.
  • Speed. Generation speed is quoted in tokens per second.

Where tokens go in a roleplay

Each request resends everything: system prompt, character card, example dialogue, lorebook entries, persona, author's note and as much history as fits. A detailed card can be 1,000 to 3,000 tokens by itself. Over a long chat, the history dominates, and every reply and every swipe pays for all of it again as input. SillyTavern shows token counts per message and a prompt breakdown (the "prompt itemization" view) so you can see what is eating the budget.

Common mistakes

  • Counting words instead of tokens when sizing cards or limits.
  • Trusting a token count from the wrong tokenizer. Counts differ between model families, sometimes by 10 to 20 percent.
  • Forgetting that output tokens, including any reasoning, are often the more expensive side.
  • Assuming images are free. In vision models, images are converted to tokens as well.
  • Estimating cost from the visible chat only. Hidden parts of the prompt, such as lorebook entries, summaries and the system prompt, are billed too, and on reasoning models the thinking is billed as output.

Tokens on Wild West API

Responses include a usage object with prompt and completion token counts (and on /v1/messages, input and output tokens), which is the authoritative number for a request. Pricing per token for each model is listed on the models page.

FAQ

How many words is 1,000 tokens?

About 750 English words, though it varies with vocabulary, names, formatting and language.

Why do input tokens grow every turn?

Chat APIs are stateless, so the frontend resends the whole conversation each time. The history gets longer, so the input does too.

Related

Uncensored AI models on one key

OpenAI and Anthropic compatible, pay as you go. New to it? Start with uncensored AI, explained.