Fireworks AI API pricing

Fireworks AI hosts open-weight models (DeepSeek, Kimi, GLM, Qwen, gpt-oss) priced per 1M tokens, with batch, priority, and fast serving options.

Official price list: docs.fireworks.ai

How Fireworks AI billing works

Serving tiers: Standard, Priority (~1.5x), Fast (~2x), and Batch at 50% of serverless — plus a +10% uplift for US-only endpoints.

Models not individually listed are priced by size class (e.g. <4B $0.10, 4–16B $0.20, >16B $0.90 per 1M tokens, flat).

Every Fireworks AI model we price

The full catalog — including dated snapshots and variants — exactly as Marginal prices incoming usage events.

307 models
Cached input / 1M
accounts/fireworks/routers/deepseek-v4p1-flash-us$0.45$1.801M$0.009
accounts/fireworks/routers/glm-5p1-fast$2.80$8.80203K$0.52
accounts/fireworks/routers/glm-5p2-fast$2.10$6.601M$0.21
accounts/fireworks/routers/glm-5p2-fast-us$2.10$6.601M$0.21
accounts/fireworks/routers/glm-5p3-fast$2.10$6.601M$0.39
accounts/fireworks/routers/glm-5p3-flash-us$0.225$0.751M$0.045
accounts/fireworks/routers/glm-5p3-us$2.10$6.601M$0.39
accounts/fireworks/routers/kimi-k2p6-fast$2.00$8.00256K$0.30
accounts/fireworks/routers/kimi-k2p7-code-fast$1.90$8.00256K$0.38
accounts/fireworks/routers/kimi-k3-fast$4.50$22.501M$0.45
accounts/fireworks/routers/kimi-k3-us$4.50$22.501M$0.45
chronos-hermes-13b-v2$0.20$0.204K—
code-llama-13b$0.20$0.2016K—
code-llama-13b-instruct$0.20$0.2016K—
code-llama-13b-python$0.20$0.2016K—
code-llama-34b$0.90$0.9016K—
code-llama-34b-instruct$0.90$0.9016K—
code-llama-34b-python$0.90$0.9016K—
code-llama-70b$0.90$0.904K—
code-llama-70b-instruct$0.90$0.904K—
code-llama-70b-python$0.90$0.904K—
code-llama-7b$0.20$0.2016K—
code-llama-7b-instruct$0.20$0.2016K—
code-llama-7b-python$0.20$0.2016K—
code-qwen-1p5-7b$0.20$0.2064K—
codegemma-2b$0.10$0.108K—
codegemma-7b$0.20$0.208K—
cogito-671b-v2-p1$1.20$1.20160K—
cogito-v1-preview-llama-3b$0.10$0.10128K—
cogito-v1-preview-llama-70b$0.90$0.90128K—
cogito-v1-preview-llama-8b$0.20$0.20128K—
cogito-v1-preview-qwen-14b$0.20$0.20128K—
cogito-v1-preview-qwen-32b$0.90$0.90128K—
dbrx-instruct$1.20$1.2032K—
deepseek-coder-1b-base$0.10$0.1016K—
deepseek-coder-33b-instruct$0.90$0.9016K—
deepseek-coder-7b-base$0.20$0.204K—
deepseek-coder-7b-base-v1p5$0.20$0.204K—
deepseek-coder-7b-instruct-v1p5$0.20$0.204K—
deepseek-coder-v2-instruct$1.20$1.2064K—
deepseek-coder-v2-lite-base$0.50$0.50160K—
deepseek-coder-v2-lite-instruct$0.50$0.50160K—
deepseek-prover-v2$1.20$1.20160K—
deepseek-r1$3.00$8.00125K—
deepseek-r1-0528$3.00$8.00160K—
deepseek-r1-0528-distill-qwen3-8b$0.20$0.20128K—
deepseek-r1-basic$0.55$2.19125K—
deepseek-r1-distill-llama-70b$0.90$0.90128K—
deepseek-r1-distill-llama-8b$0.20$0.20128K—
deepseek-r1-distill-qwen-14b$0.20$0.20128K—
deepseek-r1-distill-qwen-1p5b$0.10$0.10128K—
deepseek-r1-distill-qwen-32b$0.90$0.90128K—
deepseek-r1-distill-qwen-7b$0.20$0.20128K—
deepseek-v2-lite-chat$0.50$0.50160K—
deepseek-v2p5$1.20$1.2032K—
deepseek-v3$0.90$0.90125K—
deepseek-v3-0324$0.90$0.90160K—
deepseek-v3p1$0.56$1.68125K—
deepseek-v3p1-terminus$0.56$1.68125K—
deepseek-v3p2$0.56$1.68160K—
deepseek-v4-flash$0.14$0.281M$0.028
deepseek-v4-flash$0.14$0.281M$0.028
deepseek-v4-flash-0731$0.22$0.661M$0.007
deepseek-v4-flash-0731$0.22$0.661M$0.007
deepseek-v4-flash-vision-exp$0.22$0.661M$0.007
deepseek-v4-flash-vision-exp$0.22$0.661M$0.007
deepseek-v4-pro$1.20$1.201M$0.145
deepseek-v4-pro-0813$1.32$3.961M$0.044
deepseek-v4-pro-0813$1.32$3.961M$0.044
deepseek-v4p1-flash$0.30$1.201M$0.006
deepseek-v4p1-flash$0.30$1.201M$0.006
deepseek-v4p1-flash-us$0.45$1.801M$0.009
devstral-small-2505$0.90$0.90128K—
dobby-mini-unhinged-plus-llama-3-1-8b$0.20$0.20128K—
dobby-unhinged-llama-3-3-70b-new$0.90$0.90128K—
dolphin-2-9-2-qwen2-72b$0.90$0.90128K—
dolphin-2p6-mixtral-8x7b$0.50$0.5032K—
ember-1$3.00$15.001M$0.30
ernie-4p5-21b-a3b-pt$0.10$0.104K—
ernie-4p5-300b-a47b-pt$0.10$0.104K—
fare-20b$0.90$0.90128K—
firefunction-v1$0.50$0.5032K—
firefunction-v2$0.90$0.908K—
firellava-13b$0.20$0.204K—
firesearch-ocr-v6$0.20$0.208K—
flux-1-dev$0.10$0.104K—
flux-1-dev-controlnet-union$0.001$0.0014K—
flux-1-schnell$0.10$0.104K—
gemma-2b-it$0.10$0.108K—
gemma-3-27b-it$0.90$0.90128K—
gemma-7b$0.20$0.208K—
gemma-7b-it$0.20$0.208K—
gemma2-9b-it$0.20$0.208K—
glm-4p5$0.55$2.19125K—
glm-4p5-air$0.22$0.88125K—
glm-4p5v$1.20$1.20128K—
glm-4p6$0.55$2.19203K—
glm-4p7$0.60$2.20203K$0.30
glm-4p7$0.60$2.20203K$0.30
glm-5p1$1.40$4.40203K$0.26
glm-5p1$1.40$4.40203K$0.26
glm-5p1-fast$2.80$8.80203K$0.52
glm-5p2$1.40$4.401M$0.14
glm-5p2$1.40$4.401M$0.14
glm-5p2-fast$2.10$6.601M$0.21
glm-5p2-fast-us$2.10$6.601M$0.21
glm-5p3$1.40$4.401M$0.26
glm-5p3$1.40$4.401M$0.26
glm-5p3-fast$2.10$6.601M$0.39
glm-5p3-flash$0.15$0.501M$0.03
glm-5p3-flash$0.15$0.501M$0.03
glm-5p3-flash-us$0.225$0.751M$0.045
glm-5p3-us$2.10$6.601M$0.39
gpt-oss-120b$0.15$0.60128K$0.015
gpt-oss-120b$0.15$0.60128K$0.015
gpt-oss-20b$0.07$0.30128K$0.035
gpt-oss-20b$0.07$0.30128K$0.035
gpt-oss-safeguard-120b$1.20$1.20128K—
gpt-oss-safeguard-20b$0.50$0.50128K—
hermes-2-pro-mistral-7b$0.20$0.2032K—
inkling$1.00$4.051M$0.17
internvl3-38b$0.90$0.9016K—
internvl3-78b$0.90$0.9016K—
internvl3-8b$0.20$0.2016K—
kat-coder$0.90$0.90256K—
kat-dev-32b$0.90$0.90128K—
kat-dev-72b-exp$0.90$0.90128K—
kimi-k2-instruct$0.60$2.50128K—
kimi-k2-instruct-0905$0.60$2.50256K—
kimi-k2-thinking$0.60$2.50256K—
kimi-k2p5$0.60$3.00256K$0.10
kimi-k2p5$0.60$3.00256K$0.10
kimi-k2p6$0.95$4.00256K$0.16
kimi-k2p6$0.95$4.00256K$0.16
kimi-k2p6-fast$2.00$8.00256K$0.30
kimi-k2p7-code$0.95$4.00256K$0.19
kimi-k2p7-code$0.95$4.00256K$0.19
kimi-k2p7-code-fast$1.90$8.00256K$0.38
kimi-k3$3.00$15.001M$0.30
kimi-k3$3.00$15.001M$0.30
kimi-k3-fast$4.50$22.501M$0.45
kimi-k3-us$4.50$22.501M$0.45
llama-guard-2-8b$0.20$0.208K—
llama-guard-3-1b$0.10$0.10128K—
llama-guard-3-8b$0.20$0.20128K—
llama-v2-13b$0.20$0.204K—
llama-v2-13b-chat$0.20$0.204K—
llama-v2-70b$0.10$0.104K—
llama-v2-70b-chat$0.90$0.902K—
llama-v2-7b$0.20$0.204K—
llama-v2-7b-chat$0.20$0.204K—
llama-v3-70b-instruct$0.90$0.908K—
llama-v3-70b-instruct-hf$0.90$0.908K—
llama-v3-8b$0.20$0.208K—
llama-v3-8b-instruct-hf$0.20$0.208K—
llama-v3p1-405b-instruct$3.00$3.00125K—
llama-v3p1-405b-instruct-long$0.10$0.104K—
llama-v3p1-70b-instruct$0.90$0.90128K—
llama-v3p1-70b-instruct-1b$0.10$0.104K—
llama-v3p1-8b-instruct$0.10$0.1016K—
llama-v3p1-nemotron-70b-instruct$0.90$0.90128K—
llama-v3p2-11b-vision-instruct$0.20$0.2016K—
llama-v3p2-1b$0.10$0.10128K—
llama-v3p2-1b-instruct$0.10$0.1016K—
llama-v3p2-3b$0.10$0.10128K—
llama-v3p2-3b-instruct$0.10$0.1016K—
llama-v3p2-90b-vision-instruct$0.90$0.9016K—
llama-v3p3-70b-instruct$0.90$0.90128K—
llama4-maverick-instruct-basic$0.22$0.88128K—
llama4-scout-instruct-basic$0.15$0.60128K—
llamaguard-7b$0.20$0.204K—
llava-yi-34b$0.90$0.904K—
minimax-m1-80k$0.10$0.104K—
minimax-m2$0.30$1.204K—
minimax-m2p1$0.30$1.20200K$0.03
minimax-m2p1$0.30$1.20200K$0.03
minimax-m3$0.30$1.20500K$0.06
minimax-m3$0.30$1.20500K$0.06
ministral-3-14b-instruct-2512$0.20$0.20250K—
ministral-3-3b-instruct-2512$0.10$0.10250K—
ministral-3-8b-instruct-2512$0.20$0.20250K—
mistral-7b$0.20$0.2032K—
mistral-7b-instruct-4k$0.20$0.2032K—
mistral-7b-instruct-v0p2$0.20$0.2032K—
mistral-7b-instruct-v3$0.20$0.2032K—
mistral-7b-v0p2$0.20$0.2032K—
mistral-large-3-fp8$1.20$1.20250K—
mistral-nemo-base-2407$0.20$0.20125K—
mistral-nemo-instruct-2407$0.20$0.20125K—
mistral-small-24b-instruct-2501$0.90$0.9032K—
mixtral-8x22b$1.20$1.2064K—
mixtral-8x22b-instruct$1.20$1.2064K—
mixtral-8x22b-instruct-hf$1.20$1.2064K—
mixtral-8x7b$0.50$0.5032K—
mixtral-8x7b-instruct$0.50$0.5032K—
mixtral-8x7b-instruct-hf$0.50$0.5032K—
muse-glimmer-30b$0.35$1.50128K$0.04
muse-glimmer-30b$0.35$1.50128K$0.04
mythomax-l2-13b$0.20$0.204K—
nemotron-3-ultra-nvfp4$0.60$2.40256K$0.12
nemotron-3-ultra-nvfp4$0.60$2.40256K$0.12
nemotron-lightning-3p5-30b-a3b$0.05$0.20256K$0.01
nemotron-lightning-3p5-30b-a3b$0.05$0.20256K$0.01
nemotron-nano-v2-12b-vl$0.10$0.104K—
nous-capybara-7b-v1p9$0.20$0.2032K—
nous-hermes-2-mixtral-8x7b-dpo$0.50$0.5032K—
nous-hermes-2-yi-34b$0.90$0.904K—
nous-hermes-llama2-13b$0.20$0.204K—
nous-hermes-llama2-70b$0.90$0.904K—
nous-hermes-llama2-7b$0.20$0.204K—
nvidia-nemotron-nano-12b-v2$0.20$0.20128K—
nvidia-nemotron-nano-9b-v2$0.20$0.20128K—
openchat-3p5-0106-7b$0.20$0.208K—
openhermes-2-mistral-7b$0.20$0.2032K—
openhermes-2p5-mistral-7b$0.20$0.2032K—
openorca-7b$0.20$0.2032K—
phi-2-3b$0.10$0.102K—
phi-3-mini-128k-instruct$0.10$0.10128K—
phi-3-vision-128k-instruct$0.20$0.2032K—
phind-code-llama-34b-python-v1$0.90$0.9016K—
phind-code-llama-34b-v1$0.90$0.9016K—
phind-code-llama-34b-v2$0.90$0.9016K—
pythia-12b$0.20$0.202K—
qwen-qwq-32b-preview$0.90$0.9032K—
qwen-v2p5-14b-instruct$0.20$0.2032K—
qwen-v2p5-7b$0.20$0.20128K—
qwen1p5-72b-chat$0.90$0.9032K—
qwen2-72b-instruct$0.90$0.9032K—
qwen2-7b-instruct$0.20$0.2032K—
qwen2-vl-2b-instruct$0.10$0.1032K—
qwen2-vl-72b-instruct$0.90$0.9032K—
qwen2-vl-7b-instruct$0.20$0.2032K—
qwen2p5-0p5b-instruct$0.10$0.1032K—
qwen2p5-14b$0.20$0.20128K—
qwen2p5-1p5b-instruct$0.10$0.1032K—
qwen2p5-32b$0.90$0.90128K—
qwen2p5-32b-instruct$0.90$0.9032K—
qwen2p5-72b$0.90$0.90128K—
qwen2p5-72b-instruct$0.90$0.9032K—
qwen2p5-7b-instruct$0.20$0.2032K—
qwen2p5-coder-0p5b$0.10$0.1032K—
qwen2p5-coder-0p5b-instruct$0.10$0.1032K—
qwen2p5-coder-14b$0.20$0.2032K—
qwen2p5-coder-14b-instruct$0.20$0.2032K—
qwen2p5-coder-1p5b$0.10$0.1032K—
qwen2p5-coder-1p5b-instruct$0.10$0.1032K—
qwen2p5-coder-32b$0.90$0.9032K—
qwen2p5-coder-32b-instruct$0.90$0.904K—
qwen2p5-coder-32b-instruct-128k$0.90$0.90128K—
qwen2p5-coder-32b-instruct-32k-rope$0.90$0.9032K—
qwen2p5-coder-32b-instruct-64k$0.90$0.9064K—
qwen2p5-coder-3b$0.10$0.1032K—
qwen2p5-coder-3b-instruct$0.10$0.1032K—
qwen2p5-coder-7b$0.20$0.2032K—
qwen2p5-coder-7b-instruct$0.20$0.2032K—
qwen2p5-math-72b-instruct$0.90$0.904K—
qwen2p5-vl-32b-instruct$0.90$0.90125K—
qwen2p5-vl-3b-instruct$0.20$0.20125K—
qwen2p5-vl-72b-instruct$0.90$0.90125K—
qwen2p5-vl-7b-instruct$0.20$0.20125K—
qwen3-0p6b$0.10$0.1040K—
qwen3-14b$0.20$0.2040K—
qwen3-1p7b$0.10$0.10128K—
qwen3-1p7b-fp8-draft$0.10$0.10256K—
qwen3-1p7b-fp8-draft-131072$0.10$0.10128K—
qwen3-1p7b-fp8-draft-40960$0.10$0.1040K—
qwen3-235b-a22b$0.22$0.88128K—
qwen3-235b-a22b-instruct-2507$0.22$0.88256K—
qwen3-235b-a22b-thinking-2507$0.22$0.88256K—
qwen3-30b-a3b$0.15$0.60128K—
qwen3-30b-a3b-instruct-2507$0.50$0.50256K—
qwen3-30b-a3b-thinking-2507$0.90$0.90256K—
qwen3-32b$0.90$0.90128K—
qwen3-4b$0.20$0.2040K—
qwen3-4b-instruct-2507$0.20$0.20256K—
qwen3-8b$0.20$0.2040K—
qwen3-coder-30b-a3b-instruct$0.15$0.60256K—
qwen3-coder-480b-a35b-instruct$0.45$1.80256K—
qwen3-coder-480b-instruct-bf16$0.90$0.904K—
qwen3-next-80b-a3b-instruct$0.90$0.904K—
qwen3-next-80b-a3b-thinking$0.90$0.904K—
qwen3-vl-235b-a22b-instruct$0.22$0.88256K—
qwen3-vl-235b-a22b-thinking$0.22$0.88256K—
qwen3-vl-30b-a3b-instruct$0.15$0.60256K—
qwen3-vl-30b-a3b-thinking$0.15$0.60256K—
qwen3-vl-32b-instruct$0.90$0.904K—
qwen3-vl-8b-instruct$0.20$0.204K—
qwen3p7-plus$0.40$1.60256K$0.08
qwen3p7-plus$0.40$1.60256K$0.08
qwen3p8-max$2.00$6.00256K$0.25
qwen3p8-max$2.00$6.00256K$0.25
qwq-32b$0.90$0.90128K—
rolm-ocr$0.20$0.20125K—
snorkel-mistral-7b-pairrm-dpo$0.20$0.2032K—
stablecode-3b$0.10$0.104K—
starcoder-16b$0.20$0.208K—
starcoder-7b$0.20$0.208K—
starcoder2-15b$0.20$0.2016K—
starcoder2-3b$0.10$0.1016K—
starcoder2-7b$0.20$0.2016K—
toppy-m-7b$0.20$0.2032K—
yi-34b$0.90$0.904K—
yi-34b-200k-capybara$0.90$0.90200K—
yi-34b-chat$0.90$0.904K—
yi-6b$0.20$0.204K—
yi-large$3.00$3.0032K—
zephyr-7b-beta$0.20$0.2032K—

Prices last synced Oct 4, 2026 from the LiteLLM community price catalog — the same catalog Marginal prices real usage events against, refreshed daily. Spot an error? Tell us.

Other providers

Using Fireworks AI in production?

Marginal prices every call your app makes against this catalog and slices spend by customer, feature, or any field you define. One track() call to integrate.