Glossary
AI Token Glossary
Plain-English definitions of the key terms behind AI token calculators and LLM pricing. Use the token calculator to count tokens for your own text.
Jump to
Token
A token is the smallest unit a language model reads and generates. Rather than processing individual characters or whole words, modern LLMs break text into subword fragments and assign each fragment an integer ID. The model then operates entirely on those IDs.
In English, one token is roughly 4 characters or about three quarters of a word. Short, common words like “the”, “is”, and “in” are often a single token each. Longer or rarer words get split into two or more pieces. Punctuation, whitespace, and numbers each count separately, and code or non-Latin scripts can tokenize more expensively than English prose.
API providers bill per token for both input (what you send) and output (what the model generates). This means longer text costs more and uses more of the model's context window. Understanding tokens is the first step to predicting and controlling API costs.
Tokenizer
A tokenizer is the algorithm or program responsible for converting raw text into a sequence of tokens (and back again). It builds a fixed vocabulary of subword pieces during training, then at inference time splits every new input according to that vocabulary and maps each piece to a unique integer ID.
Every model family ships with its own tokenizer and its own vocabulary. This matters because the same text can produce different token counts depending on which model you use. A sentence that tokenizes to 50 tokens with GPT-4o might produce 47 or 55 tokens with Claude or Gemini. The difference comes from the vocabulary size, the merge rules learned during training, and how the tokenizer handles whitespace and special characters.
For English text, counts across major providers typically fall within 5 to 15 percent of each other. The gap widens for code, non-Latin scripts, and heavily formatted content. This is why token calculators aimed at non-OpenAI models often use an estimated scaling factor rather than the model's exact tokenizer.
Byte-Pair Encoding (BPE)
Byte-Pair Encoding (BPE) is a compression-inspired algorithm that builds a subword vocabulary by iteratively merging the most frequently occurring adjacent pair of units. The process starts from individual bytes or characters and applies thousands of merge operations until the vocabulary reaches a target size (for example, 50,000 or 100,000 entries).
The result is a vocabulary where common English words and word fragments each get their own token, while rare words get split into smaller pieces that still appear in the vocabulary. This balances two competing goals: a small vocabulary is efficient, but splitting every word into characters makes sequences very long and expensive. BPE sits in the middle, keeping sequences manageable while handling any word, even one that never appeared in training.
GPT models from OpenAI use BPE via their tiktoken library. Many other models also use variants of BPE, including some Llama and Mistral models. The specific merge rules and vocabulary differ across implementations, which is why identical text produces different counts on different models even when both use BPE.
SentencePiece
SentencePiece is a tokenization library developed by Google that takes a different approach from tiktoken-style BPE. Instead of requiring a language-specific pre-tokenization step (such as splitting on whitespace before applying merge rules), SentencePiece treats the entire input as a raw stream of Unicode characters, including spaces. Spaces are represented by a special underscore-like marker that you may see in tokenized output (often shown as the character “_” or a special Unicode symbol).
This design makes SentencePiece language-agnostic. It handles Japanese, Chinese, Arabic, and other scripts that do not use spaces as word boundaries just as naturally as English. It supports two main algorithms under the hood: BPE and unigram language model tokenization.
Models like Gemini, Llama, T5, and many multilingual models use SentencePiece. Because it handles whitespace differently, the same English sentence can tokenize to a slightly different count than it would with tiktoken BPE, even at a similar vocabulary size. This is one of the main reasons token counts are not interchangeable across providers.
tiktoken
tiktoken is the fast, open-source BPE tokenizer library published by OpenAI. It is what OpenAI's own API uses to count tokens for billing, so using tiktoken in a calculator produces exact counts for OpenAI models rather than estimates.
tiktoken ships with several named encodings that correspond to different model generations. The two most important are cl100k_base, which covers GPT-3.5 Turbo and GPT-4 series models, and o200k_base, which covers GPT-4o, GPT-5, and the o-series reasoning models. The number in the name refers to the approximate vocabulary size: 100,000 and 200,000 tokens respectively. The larger vocabulary means the newer encoding can represent more things in a single token, often yielding slightly lower token counts for the same text.
This token calculator runs tiktoken-compatible encodings directly in your browser using a WebAssembly build, so your text never leaves your device and counts for OpenAI models are exact.
Context Window
The context window is the maximum number of tokens a model can hold in its “working memory” at one time. It covers everything in a single request: the system prompt, the conversation history, any documents you attach, the user message, and the model's output. If the total exceeds the window, you must shorten the input or the API returns an error.
Context window sizes have grown rapidly across generations. Early GPT-3 had a 4,096-token window. Current frontier models support windows measured in hundreds of thousands or even millions of tokens, making it practical to pass entire codebases or books into a single request.
A larger context window lets you pass more documents, maintain longer conversations, and give the model more examples to work from. The trade-off is cost: more input tokens means a larger bill. There is also evidence that very long contexts can reduce the quality of attention to information in the middle of a prompt, sometimes called the “lost in the middle” problem. Keeping prompts as concise as they need to be is generally a good practice both for cost and for output quality.
Input vs Output Tokens
API providers split token billing into two buckets. Input tokens (also called prompt tokens) are everything you send to the model in a request: the system prompt, the full conversation history, any documents or context you attach, and the user's most recent message. Output tokens (also called completion tokens) are the tokens the model generates in its response.
These are priced separately because generating tokens is more computationally expensive than reading them. Output tokens typically cost several times more per million than input tokens, though the exact ratio varies by provider and model.
The formula for a request's cost is straightforward:
Total cost = (input tokens / 1,000,000) x input price per million + (output tokens / 1,000,000) x output price per million
When estimating costs for an application, you need to account for both sides. Input costs dominate for applications with long system prompts or large retrieved contexts. Output costs dominate for tasks that produce verbose responses, such as document drafting or code generation. The token calculator lets you set expected input and output lengths separately so you can model both.
Cached Input Pricing (Prompt Caching)
Prompt caching is a feature offered by several providers (including Anthropic and OpenAI) that lets you mark a portion of your prompt as cacheable. When the API sees the same cached prefix again in a later request, it charges a significantly reduced price for those tokens instead of processing them at full cost. Cache hits typically cost a fraction of normal input pricing, often around 10 to 25 percent of the standard rate, depending on the provider.
Some providers charge a small premium to write a new entry to the cache (since storing the KV state has a cost), but this write cost is usually recouped quickly if you reuse the prefix more than once or twice.
Prompt caching is most valuable when a large, static prefix appears in many requests: a long system prompt, a reference document, a code file, or a set of few-shot examples. If your application always prepends a 10,000-token document before each user question, caching that document can reduce per-request input costs substantially.
It is less useful for workloads where the prompt is unique each time or where the cacheable prefix changes frequently. Caches also expire after a period of inactivity (exact TTLs vary by provider), so very infrequent requests may not benefit.
Batch API
A Batch API is an asynchronous processing mode where you submit many requests at once (typically as a JSONL file) and receive the results later, within a promised time window that is commonly up to 24 hours. Because the provider can schedule your requests during off-peak capacity, they offer a substantial discount, often around 50 percent compared to the synchronous real-time API.
Batch processing is a good fit for workloads that do not need an instant response: evaluating a large dataset, generating descriptions for a product catalog, classifying thousands of support tickets, running a regression test suite against a new model, or processing large volumes of documents overnight.
It is a poor fit for anything user-facing that needs a real-time reply, or for pipelines where each step depends on the output of the previous one (since you may have to wait many hours between steps).
When budgeting for batch workloads, halving the per-token price can make a significant difference at scale. If you are running millions of tokens per day in non-urgent jobs, routing them through the Batch API instead of the synchronous endpoint is one of the simplest ways to cut infrastructure costs.
Token Budget
A token budget is a planned ceiling on how many tokens a given task, request, or application is allowed to use. Setting a budget serves two purposes: it keeps costs predictable and it prevents requests from inadvertently exceeding the model's context window.
A practical token budget usually covers three areas: the system prompt and static context (fixed per deployment), the dynamic context added per request (retrieved documents, conversation history, user input), and the maximum output length (controlled via the max_tokens parameter). Summing all three and checking it against the model's context limit tells you whether your design will fit.
Common techniques for staying within a budget include shortening system prompts, limiting the number of retrieved chunks in RAG pipelines, truncating or summarizing old conversation turns, and capping the max_tokens parameter to prevent runaway outputs.
Estimating your budget before building is far easier than debugging an overrun in production. Use the token calculator to measure your prompts and see exactly where your token count stands before you commit to a model or a pricing tier.
Now that you know the terminology, put it to work. Open the token calculator to count tokens and estimate costs for your own prompts, or browse the AI tools directory to find the right model for your workflow.