LLM pricing
LLM API pricing, compared
Current API prices for 59 models across OpenAI, Anthropic, Google, xAI, DeepSeek, Mistral, and the major open-weight hosts — input, output, and cached-input rates per 1M tokens, plus context windows. Updated daily from the same price catalog Marginal uses to price real usage.
| Cached input / 1M | Vendor | ||||
|---|---|---|---|---|---|
| Claude Fable 5 | $10.00 | $50.00 | 1M | $1.00 | Anthropic |
| Claude Haiku 4.5 | $1.00 | $5.00 | 200K | $0.10 | Anthropic |
| Claude Mythos 5 | $10.00 | $50.00 | 1M | $1.00 | Anthropic |
| Claude Opus 4.1legacy | $15.00 | $75.00 | 200K | $1.50 | Anthropic |
| Claude Opus 4.5legacy | $5.00 | $25.00 | 200K | $0.50 | Anthropic |
| Claude Opus 4.6legacy | $5.00 | $25.00 | 1M | $0.50 | Anthropic |
| Claude Opus 4.7legacy | $5.00 | $25.00 | 1M | $0.50 | Anthropic |
| Claude Opus 4.8legacy | $5.00 | $25.00 | 1M | $0.50 | Anthropic |
| Claude Opus 5 | $5.00 | $25.00 | 1M | $0.50 | Anthropic |
| Claude Sonnet 4.5legacy | $3.00 | $15.00 | 200K | $0.30 | Anthropic |
| Claude Sonnet 4.6 | $3.00 | $15.00 | 1M | $0.30 | Anthropic |
| Claude Sonnet 5 | $2.00 | $10.00 | 1M | $0.20 | Anthropic |
| Codestral | $0.30 | $0.90 | 32K | — | Mistral |
| DeepSeek V4 Flash | $0.44 | $1.32 | 1M | $0.014 | DeepSeek |
| DeepSeek V4 Pro | $1.32 | $3.96 | 1M | $0.044 | DeepSeek |
| Gemini 2.5 Flashlegacy | $0.30 | $2.50 | 1M | $0.03 | Google (Gemini API) |
| Gemini 2.5 Flash-Litelegacy | $0.10 | $0.40 | 1M | $0.01 | Google (Gemini API) |
| Gemini 2.5 Prolegacy | $1.25 | $10.00 | 1M | $0.125 | Google (Gemini API) |
| Gemini 3 Pro (preview)legacy | $2.00 | $12.00 | 1M | $0.20 | Google (Gemini API) |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | 1M | $0.025 | Google (Gemini API) |
| Gemini 3.1 Pro (preview) | $2.00 | $12.00 | 1M | $0.20 | Google (Gemini API) |
| Gemini 3.5 Flash | $1.50 | $9.00 | 1M | $0.15 | Google (Gemini API) |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | 1M | $0.03 | Google (Gemini API) |
| Gemini 3.6 Flash | $0.75 | $3.75 | 1M | $0.075 | Google (Gemini API) |
| Gemini 3.7 Flash | $0.75 | $3.75 | 1M | $0.075 | Google (Gemini API) |
| GLM 5.2 (Fireworks AI) | $1.40 | $4.40 | 1M | $0.14 | Z.ai |
| GPT-4.1legacy | $2.00 | $8.00 | 1M | $0.50 | OpenAI |
| GPT-4.1 minilegacy | $0.40 | $1.60 | 1M | $0.10 | OpenAI |
| GPT-4olegacy | $2.50 | $10.00 | 125K | $1.25 | OpenAI |
| GPT-4o minilegacy | $0.15 | $0.60 | 125K | $0.075 | OpenAI |
| GPT-5legacy | $1.25 | $10.00 | 272K | $0.125 | OpenAI |
| GPT-5 minilegacy | $0.25 | $2.00 | 272K | $0.025 | OpenAI |
| GPT-5 nanolegacy | $0.05 | $0.40 | 272K | $0.005 | OpenAI |
| GPT-5.1legacy | $1.25 | $10.00 | 272K | $0.125 | OpenAI |
| GPT-5.2legacy | $1.75 | $14.00 | 272K | $0.175 | OpenAI |
| GPT-5.3 Codex | $1.75 | $14.00 | 272K | $0.175 | OpenAI |
| GPT-5.4 | $2.50 | $15.00 | 1M | $0.25 | OpenAI |
| GPT-5.4 mini | $0.75 | $4.50 | 272K | $0.075 | OpenAI |
| GPT-5.4 nano | $0.20 | $1.25 | 272K | $0.02 | OpenAI |
| GPT-5.5 | $5.00 | $30.00 | 1M | $0.50 | OpenAI |
| GPT-5.6 | $5.00 | $30.00 | 1M | $0.50 | OpenAI |
| GPT-5.6 Luna | $0.20 | $1.20 | 1M | $0.02 | OpenAI |
| GPT-5.6 Terra | $2.00 | $12.00 | 1M | $0.20 | OpenAI |
| gpt-oss-120b (Groq) | $0.15 | $0.60 | 128K | $0.075 | OpenAI |
| gpt-oss-20b (Together AI) | $0.05 | $0.20 | 125K | — | OpenAI |
| Grok 4legacy | $3.00 | $15.00 | 250K | — | xAI |
| Grok 4.3 | $1.25 | $2.50 | 1M | $0.20 | xAI |
| Grok 4.5 | $2.00 | $6.00 | 500K | $0.30 | xAI |
| Grok 4.6 | $2.00 | $6.00 | 500K | $0.50 | xAI |
| Grok Build 0.1 | $1.00 | $2.00 | 250K | $0.20 | xAI |
| Kimi K2 Thinking | $0.60 | $2.50 | 256K | $0.15 | Moonshot AI |
| Kimi K2.5 | $0.60 | $3.00 | 256K | $0.10 | Moonshot AI |
| Kimi K2.6 | $0.95 | $4.00 | 256K | $0.16 | Moonshot AI |
| Ministral 3 8B | $0.15 | $0.15 | 256K | — | Mistral |
| Mistral Large 3 | $0.50 | $1.50 | 256K | — | Mistral |
| Mistral Medium 3.5 | $1.50 | $7.50 | 256K | — | Mistral |
| Mistral Small 4 | $0.15 | $0.60 | 128K | — | Mistral |
| OpenAI o3 | $2.00 | $8.00 | 200K | $0.50 | OpenAI |
| OpenAI o4-mini | $1.10 | $4.40 | 200K | $0.275 | OpenAI |
Prices last synced Aug 20, 2026 from the LiteLLM community price catalog — the same catalog Marginal prices real usage events against, refreshed daily. Spot an error? Tell us.
Pricing by provider
Every provider bills a little differently — cache multipliers, batch tiers, context-length surcharges. The provider pages list every model we price and explain the billing mechanics.
Estimate your monthly cost
Open the full calculator →Your monthly workload
Cache hit rate = share of input tokens served from the provider's prompt cache, billed at the model's cached-input price. Estimates use standard per-token rates — batch discounts, service tiers, and long-context surcharges are not applied.
| Model | Input | Cached input | Output | Per request | Monthly |
|---|---|---|---|---|---|
| Gemini 3.7 Flashcheapest | $450 | $30.00 | $938 | $0.00283 | $1,418 |
| DeepSeek V4 Pro | $792 | $17.60 | $990 | $0.00360 | $1,800 |
| Claude Sonnet 5 | $1,200 | $80.00 | $2,500 | $0.00756 | $3,780 |
| GPT-5.6 | $3,000 | $200 | $7,500 | $0.021 | $10,700 |
Popular comparisons
Frequently asked questions
How is LLM API usage billed?
Every major provider bills per token, with separate rates for input (your prompt) and output (the model's response), quoted per million tokens. Output is typically 3–6x the input rate. On top of that, most providers discount input tokens served from a prompt cache, and many offer an asynchronous batch tier at a discount.
What are cached input tokens?
When you resend the same prompt prefix (a system prompt, few-shot examples, long documents), providers can serve it from a cache instead of reprocessing it, and bill those tokens at a discounted cached-input rate — often 10x cheaper than regular input. High-traffic apps with a stable prompt prefix save the most.
Which LLM API is cheapest?
It depends on the workload — sort the table by input or output price. As a rule, each vendor's small tier (nano, mini, Flash-Lite, Haiku) is 10–100x cheaper than its flagship, and open-weight models hosted by Groq, Together, or Fireworks compete aggressively at the low end. The calculator below turns your actual request volume into a monthly figure per model.
Why does the same model have different prices on different sites?
Open-weight models (Llama, gpt-oss, Kimi, DeepSeek) are sold by many competing hosts, each setting its own price. Closed models can also be re-sold through platforms like Azure, AWS Bedrock, or OpenRouter at prices that may differ from the vendor's own API.
How up to date are these prices?
The table is generated from a price catalog that syncs daily from the LiteLLM community price map — the same catalog Marginal uses to price customers' real usage events, so errors get noticed and fixed fast. Flagship rows are additionally hand-checked against each provider's official pricing page.
List prices are the easy part.
Marginal tracks what you actually spend — every LLM call priced at that day's rates and sliced by customer, feature, or any field you define. Integration is one track() call.