Groq API pricing
Groq hosts open-weight models on custom LPU hardware, priced per 1M tokens. Its current first-party catalog centers on OpenAI's gpt-oss models and Qwen.
Official price list: console.groq.com
How Groq billing works
Prices are per 1M tokens with cached-input at roughly half the input rate; audio models bill per hour of audio instead.
The supported catalog has narrowed substantially — Llama, Kimi, and DeepSeek hostings from 2025 no longer appear in the current supported-models list. Catalog rows for them may reflect deprecated offerings.
Every Groq model we price
The full catalog — including dated snapshots and variants — exactly as Marginal prices incoming usage events.
| Cached input / 1M | ||||
|---|---|---|---|---|
| gemma-7b-it | $0.05 | $0.08 | 8K | — |
| llama-3.1-8b-instant | $0.05 | $0.08 | 128K | — |
| llama-3.3-70b-versatile | $0.59 | $0.79 | 128K | — |
| meta-llama/llama-4-maverick-17b-128e-instruct | $0.20 | $0.60 | 128K | — |
| meta-llama/llama-4-scout-17b-16e-instruct | $0.11 | $0.34 | 128K | — |
| meta-llama/llama-guard-4-12b | $0.20 | $0.20 | 8K | — |
| meta-llama/llama-prompt-guard-2-22m | $0.03 | $0.03 | 512 | — |
| meta-llama/llama-prompt-guard-2-86m | $0.04 | $0.04 | 512 | — |
| moonshotai/kimi-k2-instruct-0905 | $1.00 | $3.00 | 256K | $0.50 |
| openai/gpt-oss-120b | $0.15 | $0.60 | 128K | $0.075 |
| openai/gpt-oss-20b | $0.075 | $0.30 | 128K | $0.037 |
| openai/gpt-oss-safeguard-20b | $0.075 | $0.30 | 128K | $0.037 |
| qwen/qwen3-32b | $0.29 | $0.59 | 131K | — |
| qwen/qwen3.6-27b | $0.60 | $3.00 | 128K | — |
Prices last synced Aug 20, 2026 from the LiteLLM community price catalog — the same catalog Marginal prices real usage events against, refreshed daily. Spot an error? Tell us.
Other providers
Using Groq in production?
Marginal prices every call your app makes against this catalog and slices spend by customer, feature, or any field you define. One track() call to integrate.