Groq API pricing

Groq hosts open-weight models on custom LPU hardware, priced per 1M tokens. Its current first-party catalog centers on OpenAI's gpt-oss models and Qwen.

Official price list: console.groq.com

How Groq billing works

Prices are per 1M tokens with cached-input at roughly half the input rate; audio models bill per hour of audio instead.

The supported catalog has narrowed substantially — Llama, Kimi, and DeepSeek hostings from 2025 no longer appear in the current supported-models list. Catalog rows for them may reflect deprecated offerings.

Every Groq model we price

The full catalog — including dated snapshots and variants — exactly as Marginal prices incoming usage events.

14 models
Cached input / 1M
gemma-7b-it$0.05$0.088K
llama-3.1-8b-instant$0.05$0.08128K
llama-3.3-70b-versatile$0.59$0.79128K
meta-llama/llama-4-maverick-17b-128e-instruct$0.20$0.60128K
meta-llama/llama-4-scout-17b-16e-instruct$0.11$0.34128K
meta-llama/llama-guard-4-12b$0.20$0.208K
meta-llama/llama-prompt-guard-2-22m$0.03$0.03512
meta-llama/llama-prompt-guard-2-86m$0.04$0.04512
moonshotai/kimi-k2-instruct-0905$1.00$3.00256K$0.50
openai/gpt-oss-120b$0.15$0.60128K$0.075
openai/gpt-oss-20b$0.075$0.30128K$0.037
openai/gpt-oss-safeguard-20b$0.075$0.30128K$0.037
qwen/qwen3-32b$0.29$0.59131K
qwen/qwen3.6-27b$0.60$3.00128K

Prices last synced Aug 20, 2026 from the LiteLLM community price catalog — the same catalog Marginal prices real usage events against, refreshed daily. Spot an error? Tell us.

Other providers

Using Groq in production?

Marginal prices every call your app makes against this catalog and slices spend by customer, feature, or any field you define. One track() call to integrate.