Skip to main content

Provider pricing hub

Groq API Pricing and Models

GroqCloud models for low-latency inference and token cost planning.

Pricing summary

Groq pricing at a glance

These facts are derived from the canonical provider models currently available to CostRivo.

Models tracked

6

Cached-price models

0

Context window range

131,072 tokens

Latest verification

Jul 20, 2026

Model pricing

Groq models

Compare provider models with the same pricing unit before opening a model-specific calculator.

Swipe sideways to see all columns.

Canonical model pricing and verified model facts for this provider
ModelInput priceOutput priceCached inputContext windowCapabilitiesVerificationAction

GPT OSS 20B

GPT OSS

$0.075 / 1M tokens$0.30 / 1M tokensNot available131,072 tokenstext, reasoning
GroqGPT OSS 20B 128kStable
Official source

Verified Jul 13, 2026. Estimates vary by usage and provider pricing conditions.

Open model

Safety GPT OSS 20B

GPT OSS

$0.075 / 1M tokens$0.30 / 1M tokensNot available131,072 tokenstext, reasoning
GroqGPT OSS Safeguard 20BStable
Official source

Verified Jul 13, 2026. Estimates vary by usage and provider pricing conditions.

Open model

GPT OSS 120B

GPT OSS

$0.15 / 1M tokens$0.60 / 1M tokensNot available131,072 tokenstext, reasoning
GroqGPT OSS 120B 128kStable
Official source

Verified Jul 13, 2026. Estimates vary by usage and provider pricing conditions.

Open model

Llama 3.3 70B

Llama 3.3

$0.59 / 1M tokens$0.79 / 1M tokensNot available131,072 tokenstext
GroqLlama 3.3 70B Versatile 128kStable
Official source

Verified Jul 20, 2026. Estimates vary by usage and provider pricing conditions.

Open model

Llama 3.1 8B

Llama 3.1

$0.05 / 1M tokens$0.08 / 1M tokensNot available131,072 tokenstext
GroqLlama 3.1 8B Instant 128kStable
Official source

Verified Jul 20, 2026. Estimates vary by usage and provider pricing conditions.

Open model

Qwen/Qwen3.6-27B

Qwen 3.6

$0.60 / 1M tokens$3.00 / 1M tokensNot available131,072 tokenstext, reasoning
GroqQwen 3.6 27B 131kStable
Official source

Verified Jul 20, 2026. Estimates vary by usage and provider pricing conditions.

Open model

Price-derived model highlights

Each criterion is calculated independently from eligible, verified Groq token pricing. The same model may lead more than one criterion.

Lowest input price

Llama 3.1 8B

$0.05 / 1M

Lowest listed cost for 1M input tokens.

View model pricing

Lowest output price

Llama 3.1 8B

$0.08 / 1M

Lowest listed cost for 1M output tokens.

View model pricing

Lowest combined token price

Llama 3.1 8B

$0.13 total

Cost of 1M input tokens plus 1M output tokens.

View model pricing

Calculator

Estimate Groq API cost

Groq is preselected with GPT OSS 20B as the starting model.

Model selection

Choose the provider and model you want to estimate.
GroqGPT OSS 20B 128k

Input price

$0.075 / 1M tokens

Output price

$0.30 / 1M tokens

GroqGPT OSS 20B 128kStable
Official source

Verified Jul 13, 2026. Estimates vary by usage and provider pricing conditions.

Usage assumptions

Estimate traffic and token usage for an average request.
Active seats, customers, or internal users.
Average AI calls per user each day.
Prompt, history, and retrieved context per request.
Generated answer length; SaaS founders should test long replies.
Use 30 for always-on products or fewer for batch jobs.

Estimated results

Run the calculator to see projected cost and usage volume.

Enter your usage details, then select Calculate estimate to see your projected cost.

Estimated cost = input usage cost + output usage cost + supported optional charges.

Model catalog

Groq model catalog

Calculator actions appear only for exact model IDs with compatible verified token pricing.

8 currentReviewed Jul 15, 2026Official catalog source

8 of 8 models

Filters

Reasoning

4

GPT OSS 20B

openai/gpt-oss-20b

GPT OSS

ReasoningStableAPI available

OpenAI open-weight reasoning model hosted on GroqCloud with browser search and code execution support.

Context
131,072 tokens
Maximum output
65,536 tokens
Input
Text
Output
Text
Model owner
OpenAI
Host provider
Groq
Speed
1,000 tokens/second
Rate limits
250,000 TPM, 1,000 RPM

Capabilities

ReasoningBrowser SearchCode ExecutionFunction CallingStructured Outputs

Endpoints

Chat CompletionsResponses Compatible

Deployment

Groqcloud Api

Verified token pricing

$0.075 input / $0.3 output per 1M tokens

Use calculator
Official sourceVerified Jul 14, 2026

GPT OSS 120B

openai/gpt-oss-120b

GPT OSS

ReasoningStableAPI available

OpenAI flagship open-weight 120B model hosted on GroqCloud with reasoning and built-in tool support.

Context
131,072 tokens
Maximum output
65,536 tokens
Input
Text
Output
Text
Model owner
OpenAI
Host provider
Groq
Speed
500 tokens/second
Rate limits
250,000 TPM, 1,000 RPM

Capabilities

ReasoningBrowser SearchCode ExecutionFunction CallingStructured Outputs

Endpoints

Chat CompletionsResponses Compatible

Deployment

Groqcloud Api

Verified token pricing

$0.15 input / $0.6 output per 1M tokens

Use calculator
Official sourceVerified Jul 14, 2026

Qwen3-32B

qwen/qwen3-32b

Qwen3

ReasoningPreviewAPI available

Preview Qwen3 reasoning model hosted on GroqCloud.

Context
131,072 tokens
Maximum output
40,960 tokens
Input
Text
Output
Text
Model owner
Qwen
Host provider
Groq
Speed
400 tokens/second
Rate limits
300,000 TPM, 1,000 RPM

Capabilities

ReasoningFunction CallingStructured Outputs

Endpoints

Chat CompletionsResponses Compatible

Deployment

Groqcloud Api

Catalog details only

No compatible public calculator price is listed.

Official sourceVerified Jul 14, 2026

Qwen/Qwen3.6-27B

qwen/qwen3.6-27b

Qwen 3.6

ReasoningPreviewAPI available

Preview Qwen 3.6 27B model hosted on GroqCloud.

Context
131,072 tokens
Maximum output
32,768 tokens
Input
Text
Output
Text
Model owner
Qwen
Host provider
Groq
Speed
500 tokens/second
Rate limits
250,000 TPM, 1,000 RPM

Capabilities

ReasoningFunction CallingStructured Outputs

Endpoints

Chat CompletionsResponses Compatible

Deployment

Groqcloud Api

Verified token pricing

$0.6 input / $3 output per 1M tokens

Use calculator
Official sourceVerified Jul 14, 2026

Generative

3

Llama 4 Scout 17B 16E

meta-llama/llama-4-scout-17b-16e-instruct

Llama 4

GenerativePreviewAPI available

Preview multimodal Llama 4 Scout model hosted on GroqCloud.

Context
131,072 tokens
Maximum output
8,192 tokens
Input
Text, Image
Output
Text
Model owner
Meta
Host provider
Groq
Speed
750 tokens/second
Rate limits
300,000 TPM, 1,000 RPM

Capabilities

VisionMultimodalFunction CallingStructured Outputs

Endpoints

Chat CompletionsResponses Compatible

Deployment

Groqcloud Api

Catalog details only

No compatible public calculator price is listed.

Official sourceVerified Jul 14, 2026

Llama 3.3 70B

llama-3.3-70b-versatile

Llama 3.3

GenerativeStableAPI available

Production Llama 3.3 70B general-purpose model hosted on GroqCloud.

Context
131,072 tokens
Maximum output
32,768 tokens
Input
Text
Output
Text
Model owner
Meta
Host provider
Groq
Speed
280 tokens/second
Rate limits
300,000 TPM, 1,000 RPM

Capabilities

Text GenerationFunction CallingStructured Outputs

Endpoints

Chat CompletionsResponses Compatible

Deployment

Groqcloud Api

Verified token pricing

$0.59 input / $0.79 output per 1M tokens

Use calculator
Official sourceVerified Jul 14, 2026

Llama 3.1 8B

llama-3.1-8b-instant

Llama 3.1

GenerativeStableAPI available

Production low-latency Llama 3.1 8B model hosted on GroqCloud.

Context
131,072 tokens
Maximum output
131,072 tokens
Input
Text
Output
Text
Model owner
Meta
Host provider
Groq
Speed
560 tokens/second
Rate limits
250,000 TPM, 1,000 RPM

Capabilities

Text GenerationLow LatencyFunction Calling

Endpoints

Chat CompletionsResponses Compatible

Deployment

Groqcloud Api

Verified token pricing

$0.05 input / $0.08 output per 1M tokens

Use calculator
Official sourceVerified Jul 14, 2026

Safety

1

Safety GPT OSS 20B

openai/gpt-oss-safeguard-20b

GPT OSS

SafetyPreviewAPI available

Preview safety-focused GPT OSS model hosted on GroqCloud.

Context
131,072 tokens
Maximum output
65,536 tokens
Input
Text
Output
Text
Model owner
OpenAI
Host provider
Groq
Speed
1,000 tokens/second
Rate limits
150,000 TPM, 1,000 RPM

Capabilities

Content SafetyReasoningStructured Outputs

Endpoints

Chat CompletionsResponses Compatible

Deployment

Groqcloud Api

Verified token pricing

$0.075 input / $0.3 output per 1M tokens

Use calculator
Official sourceVerified Jul 14, 2026

Pricing sources

Pricing source and update notes

CostRivo shows official source references and verification metadata where available. Review provider pricing pages before making high-volume purchasing decisions.

Official provider pricing

Open Groq pricing references for current provider terms, tiers, and availability notes.

Open source

Pricing table

Review CostRivo's cross-provider pricing table, verification dates, source links, and lifecycle labels.

Open source

FAQ

Groq cost planning questions

Short answers for using this provider calculator.

How is Groq API cost calculated?

The calculator multiplies input, output, and cached input tokens by the selected model pricing, then scales the result by request volume.

Which Groq models have the lowest tracked token prices?

Lowest input price: Llama 3.1 8B at $0.05 per 1M tokens; Lowest output price: Llama 3.1 8B at $0.08 per 1M tokens; Lowest combined token price: Llama 3.1 8B at $0.13 for 1M input plus 1M output tokens. Each criterion is calculated independently from eligible verified token prices.

Does this estimate include cached input pricing?

Cached input pricing is not listed for the current provider models in this data set.

When was Groq pricing last verified?

The most recent model verification shown by CostRivo is Jul 20, 2026. Individual model rows retain their own source and verification details.

Pricing updates

Get pricing updates

Get notified when AI model prices change, new providers are added, product updates ship, launch notes go out, or Costrivo introduces future premium planning features.

No spam. Pricing and product updates only. We only store your email, this page, and signup time.

Optional and separate from calculator inputs.