Together AI API pricing

Together AI hosts open-weight models (Kimi, DeepSeek, GLM, Qwen, gpt-oss) priced per 1M tokens, with serverless and dedicated-endpoint options.

Official price list: www.together.ai

How Together AI billing works

Serverless is per 1M tokens with cached-input discounts on many models; dedicated endpoints bill per GPU-minute instead — high, steady traffic can cross over.

For owner-set list prices (Kimi K3, DeepSeek V4, GLM 5.2), Together's rates match Fireworks' — the differentiation is elsewhere in the stack.

Every Together AI model we price

The full catalog — including dated snapshots and variants — exactly as Marginal prices incoming usage events.

83 models
Cached input / 1M
arcee-ai/trinity-mini$0.045$0.15125K—
arize-ai/qwen-2-1.5b-instruct$0.10$0.1032K—
deepseek-ai/deepseek-coder-33b-instruct$0.80$0.8016K—
deepseek-ai/DeepSeek-R1-0528$3.00$7.00160K—
deepseek-ai/DeepSeek-R1-Distill-Llama-70B$2.00$2.00128K—
deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B$0.18$0.18128K—
deepseek-ai/DeepSeek-R1-Distill-Qwen-14B$1.60$1.60128K—
deepseek-ai/DeepSeek-V3$1.25$1.2564K—
deepseek-ai/DeepSeek-V3.1$0.60$1.70128K—
deepseek-ai/DeepSeek-V4-Flash-0731$0.14$0.281M$0.03
deepseek-ai/DeepSeek-V4-Pro-0813$1.32$3.961M$0.13
deepseek-ai/DeepSeek-V4.1-Flash$0.30$1.201M$0.006
google/gemma-2-27b-it$0.80$0.808K—
google/gemma-4-31B-it$0.39$0.97256K—
meta-llama/Llama-3-8b-chat-hf$0.20$0.208K—
meta-llama/Llama-3.1-405B-Instruct$3.50$3.504K—
meta-llama/Llama-3.2-1B-Instruct$0.06$0.06128K—
meta-llama/Llama-3.2-3B-Instruct$0.06$0.06128K—
meta-llama/Llama-3.3-70B-Instruct-Turbo$1.04$1.04128K—
meta-llama/Llama-4-Scout-17B-16E-Instruct$0.18$0.591M—
meta-llama/Meta-Llama-3-70B-Instruct-Turbo$0.88$0.888K—
meta-llama/Meta-Llama-3-8B-Instruct$0.20$0.208K—
meta-llama/Meta-Llama-3.1-70B-Instruct-Turbo$0.88$0.88128K—
meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo$0.18$0.18128K—
meta-models/Muse-Glimmer-30B$0.35$1.50128K$0.04
MiniMaxAI/MiniMax-M2.7$0.30$1.20192K$0.06
MiniMaxAI/MiniMax-M3$0.30$1.20512K$0.06
mistralai/Ministral-3-14B-Instruct-2512$0.20$0.20256K—
mistralai/Mistral-7B-Instruct-v0.1$0.20$0.2032K—
mistralai/Mistral-7B-Instruct-v0.3$0.20$0.2032K—
mistralai/Mistral-Small-24B-Instruct-2501$0.10$0.3032K—
mistralai/Mixtral-8x7B-Instruct-v0.1$0.60$0.6032K—
moonshotai/Kimi-K2-Instruct$1.00$3.00——
moonshotai/Kimi-K2.5-fp4$0.50$2.80256K—
moonshotai/Kimi-K2.6$1.20$4.50256K$0.20
moonshotai/Kimi-K2.7-Code$0.95$4.00256K$0.19
moonshotai/Kimi-K3$3.00$15.001M$0.30
NousResearch/Nous-Hermes-2-Mixtral-8x7B-DPO$0.60$0.6032K—
nvidia/Llama-3.1-Nemotron-70B-Instruct-HF$0.88$0.8832K—
nvidia/nemotron-3-ultra-550b-a55b$0.60$3.60512K$0.20
nvidia/NVIDIA-Nemotron-Nano-9B-v2$0.06$0.25128K—
openai/gpt-oss-120b$0.15$0.60128K—
openai/gpt-oss-20b$0.05$0.20128K—
Prism-ML/Ternary-Bonsai-27B$0.00$0.00256K—
Qwen/Qwen2-1.5B-Instruct$0.02$0.0232K—
Qwen/Qwen2-72B-Instruct$0.90$0.9032K—
Qwen/Qwen2-VL-72B-Instruct$1.20$1.2032K—
Qwen/Qwen2.5-14B-Instruct$0.80$0.8032K—
Qwen/Qwen2.5-72B-Instruct$1.20$1.2032K—
Qwen/Qwen2.5-72B-Instruct-Turbo$1.20$1.20128K—
Qwen/Qwen2.5-7B-Instruct-Turbo$0.30$0.3032K—
Qwen/Qwen2.5-Coder-32B-Instruct$0.80$0.8016K—
Qwen/Qwen2.5-VL-72B-Instruct$1.95$8.0032K—
Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8$2.00$2.00256K—
Qwen/Qwen3-Coder-Next-FP8$0.50$1.20256K—
Qwen/Qwen3-Next-80B-A3B-Instruct$0.15$1.50256K—
Qwen/Qwen3-Next-80B-A3B-Thinking$0.15$1.50256K—
Qwen/Qwen3-VL-32B-Instruct$0.50$1.50256K—
Qwen/Qwen3-VL-8B-Instruct$0.18$0.68256K—
Qwen/Qwen3.5-397B-A17B$0.60$3.60256K$0.35
Qwen/Qwen3.5-9B$0.17$0.25256K—
Qwen/Qwen3.6-Plus$0.50$3.001M—
Qwen/Qwen3.7-Max$1.50$4.501M$0.30
Qwen/Qwen3.7-Plus$0.32$1.281M—
Qwen/Qwen3.8-2.4T-A95B$2.00$6.001M$0.25
Qwen/Qwen3.8-Flash$0.09$0.2821M—
Qwen/QwQ-32B$1.20$1.20128K—
thinkingmachines/Inkling$1.00$4.05512K$0.17
together-ai-21.1b-41b$0.80$0.80——
together-ai-4.1b-8b$0.20$0.20——
together-ai-41.1b-80b$0.90$0.90——
together-ai-8.1b-21b$0.30$0.30——
together-ai-81.1b-110b$1.80$1.80——
together-ai-up-to-4b$0.10$0.10——
together/Tev1-4B-experimental$0.042$0.0032K$0.042
zai-org/GLM-4.5-Air-FP8$0.20$1.10128K—
zai-org/GLM-4.6$0.60$2.20198K—
zai-org/GLM-4.7$0.45$2.00198K—
zai-org/GLM-5$1.00$3.20198K—
zai-org/GLM-5.1$1.40$4.40198K$0.26
zai-org/GLM-5.2$1.40$4.401M$0.26
zai-org/GLM-5.3$1.40$4.401M$0.26
zai-org/GLM-5.3-Flash$0.15$0.501M$0.03

Prices last synced Oct 4, 2026 from the LiteLLM community price catalog — the same catalog Marginal prices real usage events against, refreshed daily. Spot an error? Tell us.

Other providers

Using Together AI in production?

Marginal prices every call your app makes against this catalog and slices spend by customer, feature, or any field you define. One track() call to integrate.