What a prompt actually costs: pricing Claude, OpenAI and Gemini API calls

What a prompt actually costs: pricing Claude, OpenAI and Gemini API calls

Every guide to LLM API pricing publishes a price table, and every price table is wrong within a month. Providers cut prices, ship new models and retire old ones faster than anyone updates a blog post.

So this is the other thing. How the pricing actually works, how to estimate what a job will cost before you run it, and the four levers that move the bill more than switching provider does. The numbers you should get from a calculator, not from an article.

How LLM API pricing works

You pay per token, quoted per million, and input and output are priced separately. Output is always the expensive one, usually several times the input rate.

A token is roughly three quarters of a word in English. A thousand words in, a thousand words out is about 2,600 tokens total. That ratio is the single most useful thing to internalise, because it turns any price sheet into an actual cost in your head.

Three things complicate that picture, and all three are worth knowing before you compare providers:

  • Cached input is charged at a fraction of the normal input rate. If you send the same long system prompt or document on every call, caching it can cut the input side dramatically.
  • Batch processing trades latency for a discount, typically around half price, on work that does not need an answer in the next second.
  • Reasoning tokens on thinking models are billed as output even though you never see most of them. A model that thinks before answering can cost several times what its headline output rate suggests.

That last one catches people out constantly. A cheap-looking reasoning model can be more expensive in practice than a pricier model that answers directly.

How to estimate a job before you run it

Four numbers, and you can do this on the back of an envelope:

  1. Input tokens per call. Your prompt plus anything attached. Words times 1.3.
  2. Output tokens per call. How long the answer will be. Words times 1.3.
  3. Calls. How many times you will run it.
  4. The two rates for the model you picked, from the provider's own pricing page.

Multiply and add. That is the whole calculation. Our token counter and cost calculator does it with live rates if you would rather paste your prompt and read the number off.

Two adjustments worth making. Add 20% for retries and failed calls, because you will have them. And if you are using a reasoning model, assume the invisible thinking is somewhere between one and five times your visible output until you have measured it on your own workload.

The cheapest LLM API is the wrong question

Ask instead which model is cheapest at an acceptable quality for your specific task, and the answer changes per task rather than per provider.

The pattern that holds across almost every workload I have costed: a small fast model handles the bulk, and a large one handles the fraction that actually needs it. Routing 80% of calls to a cheap model and 20% to an expensive one usually costs less than a fifth of running everything on the expensive one, with no visible drop in output quality.

Open-weight models hosted by an inference provider sit well below the frontier labs on price and are genuinely good enough for classification, extraction, summarisation and routing. They are usually not the right choice for work where a subtle error is expensive.

Four levers that beat switching provider

  1. Shorten the output. Output is the expensive side. Asking for 200 words instead of 800 cuts the dominant cost by three quarters and usually improves the answer.
  2. Cache the constant part. If a long system prompt or reference document rides along on every call, caching it turns your biggest input cost into a rounding error.
  3. Stop sending what you do not need. Whole documents pasted in when a section would do. Full conversation history when the last three turns would do. This is where most bloated bills come from.
  4. Batch anything that can wait. Overnight jobs, bulk classification, backfills. Around half price for work nobody is waiting on.

Do these four before you spend a week migrating to a cheaper provider. In my experience they take an afternoon and save more.

Comparing LLM pricing fairly

If you are building a comparison, three rules keep it honest.

Compare on the same task, not the same token count. A model that needs three attempts is not cheaper than one that gets it first time, whatever the per-token rate says.

Include the reasoning tokens. A comparison that counts only visible output will rank thinking models far better than they deserve.

Check the context window pricing tiers. Some providers charge more per token above a certain context length. A long-document workload can land in a higher band without you noticing.

Questions people ask

How much does the Claude API cost?

It varies by model, and Anthropic's pricing page is the only source that stays current. The structure is the same as the others: separate input and output rates per million tokens, with discounts for cached input and batch processing.

How do I count tokens in my prompt?

Roughly, words times 1.3. Exactly, use a token counter, since each provider tokenises slightly differently. Ours is free and needs no signup.

Why is my bill higher than my estimate?

Usually one of three things: reasoning tokens you were not counting, conversation history growing on every turn in a long thread, or retries. Check those before anything else.

Is the API cheaper than a subscription?

For light personal use, no. A flat monthly plan is almost always cheaper than paying per token for the same amount of chatting. The API wins when you are automating something, running it many times, or building a product.

What is the cheapest way to run a high-volume job?

A small model, batched, with the constant part of the prompt cached, and the output length capped. Those four together typically move the cost by an order of magnitude.

Where to take this next

Before optimising anything, measure one real call. Paste your actual prompt into the cost calculator, multiply by your expected volume, and see whether this is a problem worth an afternoon. Often it is not, and people spend a week saving eleven dollars.

If you are choosing between models rather than counting tokens, the LLM chooser compares them by task rather than by headline rate.

Leave a comment: