Free tool
Free AI Token Calculator
Count tokens, characters, and words for any text. Pick a provider and model, then see the estimated API cost across 80+ models including GPT, Claude, Gemini, Qwen, DeepSeek, and Llama.
Discover our library of 1,000+ free prompts and paid packs with tens of thousands more
Token visualization
Estimate the cost for a model
80% of input · ~0 output tokens
Planning usage at scale? Project your monthly and annual bill with the AI Cost Calculator.
Cost per provider*
Your input plus a 80% output, per provider. Pick any model in a card.
Anthropic
OpenAI
xAI
Z.AI
MiniMax
DeepSeek
Moonshot (Kimi)
Showing 8 of 16 providers
Model prices change weekly. Get the delta in one email.
One short email whenever a major model's pricing moves. No spam.
How the token calculator works
Large language models do not read characters or words directly. They read tokens, the small chunks of text a model is trained on. Because every API charges per token, knowing your token count is the only reliable way to predict cost before you send a request.
Paste your text above and this tool counts tokens live. OpenAI models are tokenized exactly in your browser. For Anthropic, Google, xAI, DeepSeek, Mistral, Cohere, and Perplexity, counts are estimated, since those providers do not publish an exact public tokenizer. The comparison grid shows the same text priced across the twelve most popular models so you can spot the cheapest option at a glance.
Token Ratios by Content Type
The same number of words costs more or fewer tokens depending on what the text is. These are typical ranges measured with the GPT tokenizer, so use them as estimates rather than exact figures.
| Content Type | Example | Ratio | ~Tokens / 1,000 words | Notes |
|---|---|---|---|---|
| English prose | The quick brown fox jumps over the lazy dog. | ~1.3 tokens per word | ~1,300 | Everyday writing. The GPT tokenizer is tuned for it, so this is the baseline. |
| Technical writing | The endpoint returns a 200 status on success. | ~1.5 tokens per word | ~1,500 | Jargon and uncommon terms break into more subword pieces. |
| Source code | for (let i = 0; i < n; i++) { run(i); } | ~2 to 3 tokens per word | ~2,000 to 3,000 | Brackets, operators, and indentation each tend to become their own tokens. |
| JSON or XML | { "id": 42, "active": true } | ~3 to 4 tokens per word | ~3,000 to 4,000 | Structural punctuation such as braces, quotes, and colons is token heavy. |
| Chinese, Japanese, Korean | 東京タワーは高い | ~1.5 to 2.5 tokens per character | count by character | Non-Latin scripts pack less text per token, so the same meaning costs more. |
| Numbers and IDs | 1234567890, 550e8400-e29b | varies, digits split | varies | Long numbers and identifiers fragment into several tokens each. |
How Token Pricing Works: Input, Output, Cached, and Batch
Every major LLM API charges separately for input and output tokens. Input tokens are the tokens in your prompt, system message, and any context you provide. Output tokens are what the model writes back. Because generating text is more compute-intensive than reading it, output tokens are usually priced several times higher than input tokens. On many frontier models the output rate is two to four times the input rate, so a response-heavy workload can cost far more than the raw token count suggests.
Cached input pricing is a discount providers offer when you repeatedly send the same prompt prefix, such as a long system prompt or a static document. The provider caches the key-value representations from the first call and bills subsequent calls at a reduced rate, typically 50 to 90 percent less than standard input pricing. This makes cached pricing highly valuable for applications that use a fixed context across many requests.
The batch API (offered by OpenAI, Anthropic, and others) lets you submit large numbers of requests asynchronously. Instead of real-time responses, results are returned within a set window, commonly 24 hours. In exchange for accepting this latency, providers typically charge around half the standard per-token price. Batch APIs are ideal for offline workloads such as document classification, embedding generation, or evaluation runs where immediate results are not required.
The comparison grid in the calculator above shows input, output, cached, and batch prices side by side so you can choose the pricing tier that fits your use case before committing to a model.
Who is this for?
See if your role, business, or workflow matches how people actually use this tool.
Indie AI app developers
You need to know the real per-request cost of a GPT wrapper app before shipping it, so you run an openai api cost calculator against sample prompts and responses.
Prompt engineers
You need to trim a system prompt so it stops overflowing the model's limit, so you search for a token counter for prompts to see the exact count per model.
RAG and embeddings engineers
You need to pick a chunk size that stays under the embedding model's token ceiling, so you search chunk size token calculator while splitting documents for a vector database.
Startup founders and product managers
You need to model gross margin on a new AI feature before setting a price, so you search llm api cost calculator to compare input and output token cost across providers.
AI automation agency owners and consultants
You need a defensible number to quote a client for an automation build, so you search ai project cost estimator using the client's actual prompt volume.
Customer support and chatbot builders
You need to budget a support bot handling thousands of daily conversations, so you search chatbot token cost calculator to project monthly API spend.
ML and fine-tuning engineers
You need to size a training set correctly before uploading it, so you search count tokens for fine-tuning to check a JSONL file against the provider's limits.
Content marketers and SEO teams
You need the true cost of generating hundreds of AI articles before greenlighting a batch run, so you search ai content generation cost calculator to price it out first.
Localization and multilingual product teams
You need to see how much more a non-English prompt costs to run, so you search multilingual token cost calculator and paste the same sentence in several languages.
Students and AI self-learners
You're new to how LLMs work and want a hands-on answer, so you search what is a token in ChatGPT and paste text in to watch it split apart.
No-code automation builders (n8n, Zapier, Make)
You need to know if a node's output will overflow the next step's input cap, so you search token limit calculator before wiring an AI action into your workflow.
Custom GPT and GPTs Store builders
You need your custom instructions to fit OpenAI's character cap, so you search gpt instructions character limit while trimming a system prompt down.
Technical writers and documentation teams
You need to know if an entire manual will fit in a model's context window before pasting it in, so you check a context window calculator against the doc's word count.
Finance, ops, and procurement staff
You need to forecast AI software spend for a budget review, so you search ai api budget calculator to turn expected usage into a monthly dollar figure.
AI researchers and data scientists
You need to compare context window size and price across providers before picking a model for a new project, so you search compare llm context window sizes.
Frequently asked questions
A token is the basic unit a language model reads and bills for. It is usually a word fragment of a few characters. For English text, one token is roughly four characters or about 0.75 of a word.
About 1,333 tokens for English prose (one word is roughly 1.33 tokens). See the token ratios by content type below, since code and other formats differ.
About 75 words for English prose. One token is roughly 0.75 of a word, or about four characters.
Each model family uses its own tokenizer, so counts vary by roughly 5 to 15 percent. OpenAI uses tiktoken with Byte-Pair Encoding, while Anthropic and Google use different tokenizer implementations that can produce slightly different counts for the same input.
Exact for OpenAI models (the real GPT tokenizer runs in your browser). For other providers like Anthropic, Google, and DeepSeek, counts are estimated by scaling from the GPT count, since those providers do not publish a public tokenizer. Treat estimates as close, not exact.
It depends on the model. Use the comparison grid above to compare models, or browse the full AI tools directory for detailed pricing. Costs range from under a dollar for smaller open models to several dollars per million tokens for frontier models.
The context window is the maximum number of tokens a model can consider at once, including both your input prompt and the model's generated output. Once you exceed this limit, earlier parts of the conversation are dropped.
Input tokens are the tokens you send in your prompt. Output tokens are what the model generates in response. Output tokens are usually billed at a higher rate, often two to four times the input price, because generation is more compute-intensive than processing input.
Cached input pricing is a discount applied to repeated prompt prefixes that the provider caches server-side. When you send the same system prompt repeatedly, the cached portion is billed at a reduced rate, which can significantly lower costs for high-volume applications.
Paste your prompt into the calculator above to see the token count and estimated cost for your chosen model before you send the API request. Need a prompt to test? Browse our Prompt Vault for thousands of ready-made prompts.
Vantaige has a Prompt Vault with over 1,000 free prompts, plus paid packs containing tens of thousands more, ready to paste into the calculator or your AI tool of choice.
No. Token counting runs entirely in your browser. Your text is never uploaded or stored. Only model prices are fetched, and that request contains no text.
* Prices in USD per 1,000,000 tokens, sourced from Portkey-AI/models. Prices updated 2026-09-27. (catalog: 2026-09-27) OpenAI counts are exact; other providers are estimated from the GPT tokenizer and may differ from final billing.