OpenAI cost per token
OpenAI lists API prices by the million, which is handy for billing but awkward when you are sizing one request. Divide the listed rate by 1,000,000, then multiply by your input and output token counts.
LLM API cost planning
Put in the request size once, then compare OpenAI, ChatGPT-style API, Claude, and Gemini cost estimates side by side. It is built for quick budget checks before a prompt-heavy feature goes live.
Sorted from highest to lowest using standard paid text rates. Cache writes and storage, Batch, Flex, Priority, tools, and media charges are not included.
| Model | Rates per 1M | Input | Cache read | Output | Total |
|---|---|---|---|---|---|
GPT-5.4 Pro OpenAI · 1,050K / 128K | Input $30 Cache read not listed Output $180 | $0.2400 | n/a | $0.2160 | $0.4560 |
GPT-5.5 Pro OpenAI · 1,050K / 128K | Input $30 Cache read not listed Output $180 | $0.2400 | n/a | $0.2160 | $0.4560 |
Claude Fable 5 Anthropic · 1M / 128K | Input $10 Cache read $1 Output $50 | $0.0600 | $0.002000 | $0.0600 | $0.1220 |
Claude Mythos 5 (limited availability) Anthropic · 1M / 128K | Input $10 Cache read $1 Output $50 | $0.0600 | $0.002000 | $0.0600 | $0.1220 |
GPT-5.5 OpenAI · 1,050K / 128K | Input $5 Cache read $0.5 Output $30 | $0.0300 | $0.001000 | $0.0360 | $0.0670 |
GPT-5.6 Sol OpenAI · 1,050K / 128K | Input $5 Cache read $0.5 Output $30 | $0.0300 | $0.001000 | $0.0360 | $0.0670 |
Claude Opus 4.8 Anthropic · 1M / 128K | Input $5 Cache read $0.5 Output $25 | $0.0300 | $0.001000 | $0.0300 | $0.0610 |
Claude Opus 5 Anthropic · 1M / 128K | Input $5 Cache read $0.5 Output $25 | $0.0300 | $0.001000 | $0.0300 | $0.0610 |
Claude Sonnet 4.6 Anthropic · 1M / 128K | Input $3 Cache read $0.3 Output $15 | $0.0180 | $0.000600 | $0.0180 | $0.0366 |
GPT-5.4 OpenAI · 1,050K / 128K | Input $2.5 Cache read $0.25 Output $15 | $0.0150 | $0.000500 | $0.0180 | $0.0335 |
GPT-5.6 Terra OpenAI · 1,050K / 128K | Input $2.5 Cache read $0.25 Output $15 | $0.0150 | $0.000500 | $0.0180 | $0.0335 |
Gemini 3.1 Pro Preview Google · 1,048K / 65K | Input $2 Cache read $0.2 Output $12 | $0.0120 | $0.000400 | $0.0144 | $0.0268 |
Claude Sonnet 5 Anthropic · 1M / 128K | Input $2 Cache read $0.2 Output $10 | $0.0120 | $0.000400 | $0.0120 | $0.0244 |
Gemini 3.5 Flash Google · 1,048K / 65K | Input $1.5 Cache read $0.15 Output $9 | $0.009000 | $0.000300 | $0.0108 | $0.0201 |
Gemini 2.5 Pro Google · 1,048K / 65K | Input $1.25 Cache read $0.125 Output $10 | $0.007500 | $0.000250 | $0.0120 | $0.0197 |
Gemini 3.6 Flash Google · 1,048K / 65K | Input $1.5 Cache read $0.15 Output $7.5 | $0.009000 | $0.000300 | $0.009000 | $0.0183 |
GPT-5.6 Luna OpenAI · 1,050K / 128K | Input $1 Cache read $0.1 Output $6 | $0.006000 | $0.000200 | $0.007200 | $0.0134 |
Claude Haiku 4.5 Anthropic · 200K / 64K | Input $1 Cache read $0.1 Output $5 | $0.006000 | $0.000200 | $0.006000 | $0.0122 |
GPT-5.4 Mini OpenAI · 400K / 128K | Input $0.75 Cache read $0.075 Output $4.5 | $0.004500 | $0.000150 | $0.005400 | $0.0100 |
Gemini 2.5 Flash Google · 1,048K / 65K | Input $0.3 Cache read $0.03 Output $2.5 | $0.001800 | $0.0000600 | $0.003000 | $0.004860 |
Gemini 3.5 Flash-Lite Google · 1,048K / 65K | Input $0.3 Cache read $0.03 Output $2.5 | $0.001800 | $0.0000600 | $0.003000 | $0.004860 |
Gemini 3.1 Flash-Lite Google · 1,048K / 65K | Input $0.25 Cache read $0.025 Output $1.5 | $0.001500 | $0.0000500 | $0.001800 | $0.003350 |
GPT-5.4 Nano OpenAI · 400K / 128K | Input $0.2 Cache read $0.02 Output $1.25 | $0.001200 | $0.0000400 | $0.001500 | $0.002740 |
Gemini 2.5 Flash-Lite Google · 1,048K / 65K | Input $0.1 Cache read $0.01 Output $0.4 | $0.000600 | $0.0000200 | $0.000480 | $0.001100 |
OpenAI lists API prices by the million, which is handy for billing but awkward when you are sizing one request. Divide the listed rate by 1,000,000, then multiply by your input and output token counts.
If you are building a ChatGPT-style feature on the API, count more than the visible user message. System instructions, examples, retrieved context, and the expected answer length all change the final estimate.
Claude workflows often involve long documents or long chat histories. In those cases the input side can dominate the bill, even when the generated answer is fairly short.
Gemini estimates are easiest to reason about when text input and expected output are kept separate. For images, audio, or other multimodal inputs, check Google's current pricing rules before relying on a text-only estimate.
Take the token count, divide it by 1,000,000, and multiply by the provider's listed rate. For a real request, do that separately for input and output, then add the two numbers.
Often, yes. Many providers price generated tokens higher than prompt tokens, so a short prompt with a long answer can still become the expensive part of a request.
Cached input pricing is a lower rate for repeated prompt content, such as a stable system prompt or fixed document prefix. Only count text as cached when the provider would actually treat it that way.
No. Treat it as a planning tool before a request. For production numbers, use the provider's current pricing page and your billing dashboard.