Together AI API pricing
Together AI hosts open-weight models (Kimi, DeepSeek, GLM, Qwen, gpt-oss) priced per 1M tokens, with serverless and dedicated-endpoint options.
Official price list: www.together.ai
How Together AI billing works
Serverless is per 1M tokens with cached-input discounts on many models; dedicated endpoints bill per GPU-minute instead — high, steady traffic can cross over.
For owner-set list prices (Kimi K3, DeepSeek V4, GLM 5.2), Together's rates match Fireworks' — the differentiation is elsewhere in the stack.
Every Together AI model we price
The full catalog — including dated snapshots and variants — exactly as Marginal prices incoming usage events.
| Cached input / 1M | ||||
|---|---|---|---|---|
| deepseek-ai/DeepSeek-R1 | $3.00 | $7.00 | 125K | — |
| deepseek-ai/DeepSeek-R1-0528-tput | $0.55 | $2.19 | 125K | — |
| deepseek-ai/DeepSeek-V3 | $1.25 | $1.25 | 64K | — |
| deepseek-ai/DeepSeek-V3.1 | $0.60 | $1.70 | 125K | — |
| meta-llama/Llama-3.3-70B-Instruct-Turbo | $0.88 | $0.88 | — | — |
| meta-llama/Llama-3.3-70B-Instruct-Turbo-Free | $0.00 | $0.00 | — | — |
| meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8 | $0.27 | $0.85 | — | — |
| meta-llama/Llama-4-Scout-17B-16E-Instruct | $0.18 | $0.59 | — | — |
| meta-llama/Meta-Llama-3.1-405B-Instruct-Turbo | $3.50 | $3.50 | — | — |
| meta-llama/Meta-Llama-3.1-70B-Instruct-Turbo | $0.88 | $0.88 | — | — |
| meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo | $0.18 | $0.18 | — | — |
| mistralai/Mixtral-8x7B-Instruct-v0.1 | $0.60 | $0.60 | — | — |
| moonshotai/Kimi-K2-Instruct | $1.00 | $3.00 | — | — |
| moonshotai/Kimi-K2-Instruct-0905 | $1.00 | $3.00 | 256K | — |
| moonshotai/Kimi-K2.5 | $0.50 | $2.80 | 250K | — |
| openai/gpt-oss-120b | $0.15 | $0.60 | 128K | — |
| openai/gpt-oss-20b | $0.05 | $0.20 | 125K | — |
| Qwen/Qwen3-235B-A22B-fp8-tput | $0.20 | $0.60 | 40K | — |
| Qwen/Qwen3-235B-A22B-Instruct-2507-tput | $0.20 | $6.00 | 262K | — |
| Qwen/Qwen3-235B-A22B-Thinking-2507 | $0.65 | $3.00 | 250K | — |
| Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8 | $2.00 | $2.00 | 250K | — |
| Qwen/Qwen3-Next-80B-A3B-Instruct | $0.15 | $1.50 | 256K | — |
| Qwen/Qwen3-Next-80B-A3B-Thinking | $0.15 | $1.50 | 256K | — |
| Qwen/Qwen3.5-397B-A17B | $0.60 | $3.60 | 256K | — |
| together-ai-21.1b-41b | $0.80 | $0.80 | — | — |
| together-ai-4.1b-8b | $0.20 | $0.20 | — | — |
| together-ai-41.1b-80b | $0.90 | $0.90 | — | — |
| together-ai-8.1b-21b | $0.30 | $0.30 | — | — |
| together-ai-81.1b-110b | $1.80 | $1.80 | — | — |
| together-ai-up-to-4b | $0.10 | $0.10 | — | — |
| zai-org/GLM-4.5-Air-FP8 | $0.20 | $1.10 | 125K | — |
| zai-org/GLM-4.6 | $0.60 | $2.20 | 200K | — |
| zai-org/GLM-4.7 | $0.45 | $2.00 | 200K | — |
Prices last synced Aug 20, 2026 from the LiteLLM community price catalog — the same catalog Marginal prices real usage events against, refreshed daily. Spot an error? Tell us.
Other providers
Using Together AI in production?
Marginal prices every call your app makes against this catalog and slices spend by customer, feature, or any field you define. One track() call to integrate.