LLM pricing

LLM API pricing, compared

Current API prices for 59 models across OpenAI, Anthropic, Google, xAI, DeepSeek, Mistral, and the major open-weight hosts — input, output, and cached-input rates per 1M tokens, plus context windows. Updated daily from the same price catalog Marginal uses to price real usage.

59 models
Cached input / 1MVendor
Claude Fable 5$10.00$50.001M$1.00Anthropic
Claude Haiku 4.5$1.00$5.00200K$0.10Anthropic
Claude Mythos 5$10.00$50.001M$1.00Anthropic
Claude Opus 4.1legacy$15.00$75.00200K$1.50Anthropic
Claude Opus 4.5legacy$5.00$25.00200K$0.50Anthropic
Claude Opus 4.6legacy$5.00$25.001M$0.50Anthropic
Claude Opus 4.7legacy$5.00$25.001M$0.50Anthropic
Claude Opus 4.8legacy$5.00$25.001M$0.50Anthropic
Claude Opus 5$5.00$25.001M$0.50Anthropic
Claude Sonnet 4.5legacy$3.00$15.00200K$0.30Anthropic
Claude Sonnet 4.6$3.00$15.001M$0.30Anthropic
Claude Sonnet 5$2.00$10.001M$0.20Anthropic
Codestral$0.30$0.9032KMistral
DeepSeek V4 Flash$0.44$1.321M$0.014DeepSeek
DeepSeek V4 Pro$1.32$3.961M$0.044DeepSeek
Gemini 2.5 Flashlegacy$0.30$2.501M$0.03Google (Gemini API)
Gemini 2.5 Flash-Litelegacy$0.10$0.401M$0.01Google (Gemini API)
Gemini 2.5 Prolegacy$1.25$10.001M$0.125Google (Gemini API)
Gemini 3 Pro (preview)legacy$2.00$12.001M$0.20Google (Gemini API)
Gemini 3.1 Flash-Lite$0.25$1.501M$0.025Google (Gemini API)
Gemini 3.1 Pro (preview)$2.00$12.001M$0.20Google (Gemini API)
Gemini 3.5 Flash$1.50$9.001M$0.15Google (Gemini API)
Gemini 3.5 Flash-Lite$0.30$2.501M$0.03Google (Gemini API)
Gemini 3.6 Flash$0.75$3.751M$0.075Google (Gemini API)
Gemini 3.7 Flash$0.75$3.751M$0.075Google (Gemini API)
GLM 5.2 (Fireworks AI)$1.40$4.401M$0.14Z.ai
GPT-4.1legacy$2.00$8.001M$0.50OpenAI
GPT-4.1 minilegacy$0.40$1.601M$0.10OpenAI
GPT-4olegacy$2.50$10.00125K$1.25OpenAI
GPT-4o minilegacy$0.15$0.60125K$0.075OpenAI
GPT-5legacy$1.25$10.00272K$0.125OpenAI
GPT-5 minilegacy$0.25$2.00272K$0.025OpenAI
GPT-5 nanolegacy$0.05$0.40272K$0.005OpenAI
GPT-5.1legacy$1.25$10.00272K$0.125OpenAI
GPT-5.2legacy$1.75$14.00272K$0.175OpenAI
GPT-5.3 Codex$1.75$14.00272K$0.175OpenAI
GPT-5.4$2.50$15.001M$0.25OpenAI
GPT-5.4 mini$0.75$4.50272K$0.075OpenAI
GPT-5.4 nano$0.20$1.25272K$0.02OpenAI
GPT-5.5$5.00$30.001M$0.50OpenAI
GPT-5.6$5.00$30.001M$0.50OpenAI
GPT-5.6 Luna$0.20$1.201M$0.02OpenAI
GPT-5.6 Terra$2.00$12.001M$0.20OpenAI
gpt-oss-120b (Groq)$0.15$0.60128K$0.075OpenAI
gpt-oss-20b (Together AI)$0.05$0.20125KOpenAI
Grok 4legacy$3.00$15.00250KxAI
Grok 4.3$1.25$2.501M$0.20xAI
Grok 4.5$2.00$6.00500K$0.30xAI
Grok 4.6$2.00$6.00500K$0.50xAI
Grok Build 0.1$1.00$2.00250K$0.20xAI
Kimi K2 Thinking$0.60$2.50256K$0.15Moonshot AI
Kimi K2.5$0.60$3.00256K$0.10Moonshot AI
Kimi K2.6$0.95$4.00256K$0.16Moonshot AI
Ministral 3 8B$0.15$0.15256KMistral
Mistral Large 3$0.50$1.50256KMistral
Mistral Medium 3.5$1.50$7.50256KMistral
Mistral Small 4$0.15$0.60128KMistral
OpenAI o3$2.00$8.00200K$0.50OpenAI
OpenAI o4-mini$1.10$4.40200K$0.275OpenAI

Prices last synced Aug 20, 2026 from the LiteLLM community price catalog — the same catalog Marginal prices real usage events against, refreshed daily. Spot an error? Tell us.

Pricing by provider

Every provider bills a little differently — cache multipliers, batch tiers, context-length surcharges. The provider pages list every model we price and explain the billing mechanics.

Estimate your monthly cost

Open the full calculator →

Your monthly workload

Cache hit rate = share of input tokens served from the provider's prompt cache, billed at the model's cached-input price. Estimates use standard per-token rates — batch discounts, service tiers, and long-context surcharges are not applied.

GPT-5.6Claude Sonnet 5Gemini 3.7 FlashDeepSeek V4 Pro
ModelInputCached inputOutputPer requestMonthly
Gemini 3.7 Flashcheapest$450$30.00$938$0.00283$1,418
DeepSeek V4 Pro$792$17.60$990$0.00360$1,800
Claude Sonnet 5$1,200$80.00$2,500$0.00756$3,780
GPT-5.6$3,000$200$7,500$0.021$10,700
Rates used: GPT-5.6 $5.00 in / $30.00 out · Claude Sonnet 5 $2.00 in / $10.00 out · Gemini 3.7 Flash $0.75 in / $3.75 out · DeepSeek V4 Pro $1.32 in / $3.96 out

Popular comparisons

Frequently asked questions

How is LLM API usage billed?

Every major provider bills per token, with separate rates for input (your prompt) and output (the model's response), quoted per million tokens. Output is typically 3–6x the input rate. On top of that, most providers discount input tokens served from a prompt cache, and many offer an asynchronous batch tier at a discount.

What are cached input tokens?

When you resend the same prompt prefix (a system prompt, few-shot examples, long documents), providers can serve it from a cache instead of reprocessing it, and bill those tokens at a discounted cached-input rate — often 10x cheaper than regular input. High-traffic apps with a stable prompt prefix save the most.

Which LLM API is cheapest?

It depends on the workload — sort the table by input or output price. As a rule, each vendor's small tier (nano, mini, Flash-Lite, Haiku) is 10–100x cheaper than its flagship, and open-weight models hosted by Groq, Together, or Fireworks compete aggressively at the low end. The calculator below turns your actual request volume into a monthly figure per model.

Why does the same model have different prices on different sites?

Open-weight models (Llama, gpt-oss, Kimi, DeepSeek) are sold by many competing hosts, each setting its own price. Closed models can also be re-sold through platforms like Azure, AWS Bedrock, or OpenRouter at prices that may differ from the vendor's own API.

How up to date are these prices?

The table is generated from a price catalog that syncs daily from the LiteLLM community price map — the same catalog Marginal uses to price customers' real usage events, so errors get noticed and fixed fast. Flagship rows are additionally hand-checked against each provider's official pricing page.

List prices are the easy part.

Marginal tracks what you actually spend — every LLM call priced at that day's rates and sliced by customer, feature, or any field you define. Integration is one track() call.