Groq Pricing (2026): Throughput, Not Tokens, Forces the Upgrade

Groq pricing verified August 2026: free and paid limits, every model rate, batch and cache rules, hidden tool costs, and provider comparisons.

Saturday, August 15, 2026Omid Saffari
Tools
  • GGroq
  • TTogether AI
  • FFireworks AI
Groq Pricing (2026): Throughput, Not Tokens, Forces the Upgrade

Groq costs $0 until its model-specific Free limits stop your workflow; after that, Developer has no subscription fee and charges only for usage. On a worked GPT-OSS 120B load, the Free daily token ceiling represents just $1.44 a month of paid traffic, so throughput, not token spend, is what forces the upgrade.

Groq pricing at a glance

Groq has no self-serve monthly seat price to memorize. The choice is between a rate-limited Free plan, a pay-as-you-go Developer plan, and custom Enterprise capacity.

These prices and limits were verified against Groq's live pricing page, Supported Models, Billing FAQs, and rate-limit documentation on August 15, 2026.

Groq live pricing page with model and built-in tool rates
Groq pricing, verified August 2026
TierPriceCapacityBest fit
Free$0Model-specific organization limitsEvaluation and light internal use
DeveloperNo base subscription; usage billedBase text-model limits up to 300K TPM and 1K RPMSelf-serve production
EnterpriseCustom provisioned capacityContracted input and output bundlesCritical paths needing an SLA

A token is the unit the model bills. Input tokens include instructions, conversation history, documents, and tool definitions sent to the model. Output tokens are what the model generates. Groq quotes both sides per one million tokens, so monthly cost is the sum of input, output, audio, and any built-in tool usage.

Developer does not carry a monthly or annual subscription commitment. Groq asks for a payment method, meters usage, and normally invoices in arrears. There is no separate overage rate: Free blocks excess traffic at the applicable limit, while Developer continues billing usage at the same model rates until an account or Spend Limit blocks it. Enterprise Performance works differently: the buyer purchases provisioned input and output capacity rather than paying the public on-demand token rate.

The Free plan is useful until a shared limit becomes a production dependency

Groq's Free plan can support a prototype, an internal helper, and a surprisingly useful amount of low-volume automation. It stops being free infrastructure the moment a rate-limit error becomes a customer-facing event.

Limits apply to the organization, not to each API key. Creating more keys does not multiply capacity. Groq also enforces several dimensions at once: requests per minute, requests per day, tokens per minute, and tokens per day. The first ceiling reached blocks the next call with HTTP 429.

For the current production text models, the published Free limits include:

  • GPT-OSS 120B and GPT-OSS 20B: 30 RPM, 1,000 RPD, 8,000 TPM, and 200,000 TPD.
  • Llama 3.3 70B Versatile: 30 RPM, 1,000 RPD, 12,000 TPM, and 100,000 TPD.
  • Llama 3.1 8B Instant: 30 RPM, 14,400 RPD, 6,000 TPM, and 500,000 TPD.
  • Compound and Compound Mini: 30 RPM, 250 RPD, and 70,000 TPM.
  • Whisper Large V3 and Turbo: 20 RPM, 2,000 RPD, 7,200 audio seconds per hour, and 28,800 audio seconds per day.

Groq describes that public table as a high-level summary. The Limits page inside the Console is the authority for an individual organization's current allowance.

The daily token ceiling often binds before the request count. Take a GPT-OSS 120B request with 2,000 input tokens and 500 output tokens. At 2,500 total tokens, the 200,000 TPD limit permits 80 such requests a day. The 8,000 TPM ceiling permits three in one minute; a fourth comparable request would cross the token limit even though the plan advertises 30 RPM.

That is the free-to-paid break-even. The same request costs $0.0006 on Developer, so 80 a day for a 30-day month costs about $1.44. There is little financial reason to bend a production workflow around the Free ceiling. The upgrade buys capacity while the model bill is still smaller than most SaaS add-ons.

Decision flow from Groq Free through Developer to Enterprise for a worked GPT-OSS 120B workload
For the worked 2,500-token request, Free reaches 80 per day while the same monthly traffic costs about $1.44 on Developer

Who never needs to pay

You may never need Developer if Groq is used for local development, occasional evaluation, a personal internal script, or a demo whose traffic remains below the applicable token and request caps. Free also makes sense when a 429 can wait for a retry and nobody is promised continuous service.

Pay when the workflow has a user, a deadline, or a queue that cannot tolerate the shared cap. Developer also adds Batch, Flex, chat support, and Spend Limits. Those operational controls, not a different model catalog, are the practical boundary.

Every current Groq model price

Groq's current production catalog is compact: four text models, two speech-to-text models, and two Compound systems. A separate preview catalog carries models that may disappear on short notice.

Production text models

  • Llama 3.1 8B Instant costs $0.05 input and $0.08 output per million tokens. It is the cheapest production text route. Its base Developer allowance is 250,000 TPM and 1,000 RPM, with a 131,072-token context window.
  • GPT-OSS 20B costs $0.075 input and $0.30 output per million tokens. Cached input is $0.0375 per million. Its base Developer allowance is 250,000 TPM and 1,000 RPM, with a 131,072-token context window.
  • GPT-OSS 120B costs $0.15 input and $0.60 output per million tokens. Cached input is $0.075 per million. Its base Developer allowance is 250,000 TPM and 1,000 RPM, with a 131,072-token context window.
  • Llama 3.3 70B Versatile costs $0.59 input and $0.79 output per million tokens. Its base Developer allowance is 300,000 TPM and 1,000 RPM, with a 131,072-token context window.

Start a production evaluation with GPT-OSS 120B when a larger open-weight model is required but a same-model provider fallback matters. Move down to GPT-OSS 20B or Llama 3.1 8B only when an acceptance check shows the smaller model keeps the result quality. Move to Llama 3.3 70B only when its outputs earn a rate that is 3.93 times GPT-OSS 120B on input and 1.32 times on output.

Those multiples do not prove one model is better. They define what the outcome must repay through fewer retries, shorter responses, lower review work, or a capability the cheaper route lacks.

Speech-to-text models

  • Whisper Large V3 costs $0.111 per audio hour. Its base Developer capacity is 200,000 audio seconds per hour and 300 RPM, with a 100 MB maximum file.
  • Whisper Large V3 Turbo costs $0.04 per audio hour. Its base Developer capacity is 400,000 audio seconds per hour and 400 RPM.

At 1,000 audio hours, the line items are $111 for Whisper Large V3 and $40 for Turbo, a $71 gap. The lower price is the starting call, but an audio team should keep whichever model clears its transcript acceptance threshold. A cheaper transcript that needs correction is not a cheaper outcome.

Compound systems and preview models

Groq Compound and Compound Mini combine production models with server-side tools. They do not have one flat token rate. Groq passes through the underlying model and tool charges, which is why a Compound bill needs its own cost stack rather than a single per-token estimate.

The live preview catalog currently lists:

  • Qwen3.6 27B: $0.60 input and $3.00 output per million tokens.
  • Safety GPT-OSS 20B: $0.075 input and $0.30 output per million tokens.
  • Llama Prompt Guard 2 22M: $0.03 input and output per million tokens.
  • Llama Prompt Guard 2 86M: $0.04 input and output per million tokens.
  • Canopy Labs Orpheus V1 English: $22 per million characters.
  • Canopy Labs Orpheus Arabic Saudi: $40 per million characters.
  • MiniMax M2.7: Enterprise pricing by sales inquiry.

Preview means evaluation, not a production promise. Groq says these models can be discontinued at short notice. A budget that depends on Qwen3.6 27B therefore also needs a tested route to another provider or model before it reaches customers.

Groq's current Supported Models page does not list a DeepSeek or Kimi model. Older posts and calculators that still price those names as active Groq routes are stale. If DeepSeek is the requirement, price it through a current provider rather than assuming it remains in Groq's catalog; the DeepSeek pricing analysis covers that separate decision.

Cost per outcome decides the model

Groq's token rates are so low that model choice matters less than traffic shape until usage becomes large. The clean comparison fixes one request and prices each route against it.

Use a standard worked request with 2,000 input tokens and 500 output tokens. At 100,000 requests a month, before caching, Batch, tools, or retries:

  • Llama 3.1 8B Instant: $14 total, or $0.00014 per request.
  • GPT-OSS 20B: $30 total, or $0.0003 per request.
  • GPT-OSS 120B: $60 total, or $0.0006 per request.
  • Llama 3.3 70B Versatile: $157.50 total, or $0.001575 per request.
  • Qwen3.6 27B preview: $270 total, or $0.0027 per request.

The useful denominator is not a request. It is an accepted result. If a smaller model retries, produces longer output, or sends more work to a human reviewer, the cheap request can become the expensive workflow. Track accepted outcomes beside tokens so the model can earn its place instead of winning on sticker price.

There is no paid volume break-even where one Groq model automatically becomes cheaper than another. Every public on-demand rate is linear and there is no base fee. GPT-OSS 20B stays half the worked cost of GPT-OSS 120B at one request and at one million requests. The decision flips only when outcome quality or operating behavior changes the amount of work needed.

For a funded founder, that can mean GPT-OSS 20B for classification and GPT-OSS 120B for ambiguous cases. For a mid-market CTO, it means retaining model and token fields in every request log so procurement can see cost per accepted workflow. For a senior operator, it means measuring corrections and retries, not treating one successful demo as proof of the production route.

Batch beats caching for asynchronous work, and the discounts do not stack

Groq Batch is the lowest published rate when a result can wait. It charges 50% less than synchronous API pricing, runs outside standard per-model limits, and accepts a processing window from 24 hours to 7 days.

On the 100,000-request GPT-OSS 120B workload, the choices are:

  • Synchronous with no cache: $60.
  • Synchronous with 50% of input tokens served from cache: $52.50.
  • Synchronous with 80% of input tokens served from cache: $48.
  • Batch: $30.

The current rule that changes the budget is explicit: Batch and prompt-caching discounts do not stack. Every Batch token is billed at the 50% Batch rate regardless of cache status. A spreadsheet that halves the $30 Batch total again to $15 is wrong.

Prompt caching remains useful for synchronous GPT-OSS traffic. It is automatic, carries no feature fee, and discounts cached input by 50%. It currently supports GPT-OSS 20B, GPT-OSS 120B, and GPT-OSS Safeguard 20B. Cache hits require an exact matching prefix, the minimum cacheable prefix ranges from 128 to 1,024 tokens by model, and unused cache data expires after two hours. A hit is best effort, not guaranteed.

Put stable system instructions and tool definitions at the beginning of a prompt, with variable user data after them. Then read cached-token usage from responses. A projected cache ratio is not a budget fact until production usage shows it.

Batch has different operational edges. A JSONL file can contain up to 50,000 lines and be up to 200 MB. Jobs that miss the chosen processing window expire, although Groq charges only successfully completed requests. Batch files and results can remain in Groq's systems for up to 30 days, so sensitive workloads need a data-retention review before cost savings decide the route.

$60 of tokens can become a $560 Compound bill

Compound's hidden cost is not the model. It is the tool loop.

Groq prices Basic Search at $5 per 1,000 requests, Advanced Search at $8 per 1,000, Visit Website at $1 per 1,000, and Code Execution at $0.18 per hour. Those charges sit on top of the underlying model tokens.

Take 100,000 standard GPT-OSS 120B requests. The model tokens cost $60. If every request makes one Basic Search call, search adds $500 and the combined bill becomes $560. One Advanced Search each produces an $860 total. One Visit Website each produces $160.

Cost columns comparing 60 dollars of GPT-OSS tokens with 500 dollars of Basic Search charges
At 100,000 worked requests, one Basic Search each adds $500 to a $60 token bill

This is the budget line most likely to surprise an agent team. Token optimization cannot rescue a workflow whose agent searches on every turn. The first control is a tool policy: define when search is necessary, cap tool loops, cache durable results outside the model, and record tool calls per accepted outcome.

The paid bill has no annual lock, but it has operational edges

Developer is usage-billed, not a disguised annual SaaS contract. The less obvious terms concern when money leaves the account, how limits behave, and what happens to credits.

New Developer accounts use progressive billing. Groq can issue an invoice when cumulative lifetime usage crosses $1, $10, $100, $500, and $1,000. After the $1,000 threshold, ordinary monthly billing takes over. Accounts with an Indian billing address use $1 and $10 thresholds followed by recurring $100 increments. Groq does not require payment until usage reaches at least $0.50.

That schedule can create several small card charges early in adoption. It is not an added fee, but finance should recognize the pattern before treating the charges as duplicates.

Spend Limits block traffic, with a short reporting lag

Paid plans can set one organization-wide monthly Spend Limit. It covers every API key and endpoint, resets on the first of the month, and blocks new requests once the limit is detected. Groq supports alert points such as 50%, 75%, and 90% of the cap.

Spend tracking updates with a 10 to 15 minute delay. A burst can therefore push the final bill slightly past the configured amount before blocking starts. Groq suggests beginning with a limit 20% to 30% above expected monthly usage. For a critical path, that buffer also reduces the chance that a normal traffic spike turns the budget control into an outage.

Credits expire, and prepaid sales are final

Groq Service Credits expire one year after purchase or issuance unless a different term is specified. They are non-transferable, cannot be redeemed for cash, and are non-refundable. Closing the account expires the balance immediately.

Groq's Billing FAQ handles general refund requests case by case through support. That does not override the specific Service Credit terms: purchased and promotional credits remain final-sale balances. Buy a credit amount tied to a forecast, not a vague future migration.

Flex trades guaranteed capacity for a higher ceiling

Flex is available to paid customers at the same token price as On Demand and offers 10 times the on-demand rate limit. The trade is best-effort capacity. If Flex is full, it can fail quickly with HTTP 498 and capacity_exceeded.

Use Flex for retryable queues, evaluations, and bulk work that can absorb a fast failure. Do not make it the only route for a non-repeatable customer action. On Demand has predictable speed but can add queue latency during peaks. Enterprise Performance is the contracted option when that variance is unacceptable, with a 99.9% availability SLA and a 99% latency guarantee under the enterprise agreement.

Groq, Together AI, and Fireworks tie on uncached GPT-OSS 120B tokens

Groq does not win GPT-OSS 120B on public token price alone. Groq, Together AI, and Fireworks AI all list $0.15 input and $0.60 output per million tokens for the same model.

Use one normalized workload: 100 million input tokens and 20 million output tokens. With no cache discount, each provider costs $27. That tie is the useful result because it forces the provider decision onto cache economics, rate limits, deployment choices, reliability, and observed latency.

Together AI

Together AI is a multi-model inference provider whose serverless GPT-OSS 120B route costs $0.15 input and $0.60 output per million tokens. Its serverless product has no minimum or provisioning charge.

Together AI live serverless pricing page showing GPT-OSS 120B rates
Together AI GPT-OSS 120B pricing

Together's current GPT-OSS 120B row does not publish a cached-input rate. Its documentation says a model without a cached price bills all input at the standard rate. On the normalized workload, that leaves the total at $27 even if the prompt repeats. Together offers up to 50% lower Batch pricing on supported serverless models, so it remains viable for asynchronous work when that exact model is eligible.

Choose Together when the broader model catalog or a path from serverless to reserved infrastructure matters more than Groq's specific operating profile. Skip a price-led move between the two until a workload exposes a difference, because the uncached GPT-OSS line item is identical.

Fireworks AI

Fireworks AI lists GPT-OSS 120B at $0.15 input, $0.014 cached input, and $0.60 output per million tokens, with $1 of serverless signup credit.

Fireworks AI GPT-OSS 120B model page with standard and cached token prices
Fireworks AI GPT-OSS 120B pricing

At a 50% input cache share, the normalized total is $20.20 on Fireworks, $23.25 on Groq, and $27 on Together. That makes Fireworks the public-price winner for repeat-heavy GPT-OSS 120B input, provided its listed cached rate applies to the measured workload. Groq Batch is still lower for asynchronous work at $13.50 on the same normalized traffic.

OpenAI developed GPT-OSS but does not serve it through the OpenAI API or its consumer app. A direct OpenAI API comparison therefore swaps both provider and model, which breaks price normalization. Compare the same model across hosting providers first, then evaluate different proprietary models by accepted outcome. The cheapest AI API shortlist covers the broader model and provider field.

The provider decision is straightforward:

  • Choose Groq when synchronous throughput and its operating controls win the measured route, or when 50% Batch pricing makes delayed work cheapest.
  • Choose Fireworks when repeated GPT-OSS context consistently earns the listed $0.014 cached-input rate.
  • Choose Together when its wider serverless-to-dedicated path matters and the workload does not benefit from a GPT-OSS cached rate.

Who should pick Groq, and who should skip it

Groq is a strong choice when an open model already clears the quality bar and response speed affects the product. It is a weaker choice when the model requirement, capacity guarantee, or deployment control sits outside its current self-serve catalog.

The upside
What it does well
6 points

  • The Free plan is useful enough to evaluate production models before adding a payment method.
  • Developer has no base subscription or annual lock.
  • Public on-demand rates remain below $1 per million tokens for every production text model.
  • Batch cuts eligible asynchronous work by 50% without consuming standard model limits.
  • GPT-OSS prompt caching is automatic and halves cached input rates.
  • Spend Limits can block organization-wide usage before a monthly budget runs unchecked.
The downside
Where it falls short
6 points

  • Free limits are shared at the organization level and can bind on tokens long before requests.
  • Batch and caching discounts do not stack.
  • Compound tool charges can exceed model-token spend by a wide margin.
  • Prompt caching supports only the listed GPT-OSS family and depends on exact prefix reuse.
  • Flex can return HTTP 498 when best-effort capacity is unavailable.
  • Preview models can be removed at short notice, so they cannot carry an unportable production dependency.

A funded founder should choose Groq when a fast open model keeps the product responsive and the team can preserve a same-model fallback. A mid-market CTO should use Developer only after centralizing keys, logging model and tool usage, and setting an organization cap. A senior operator should care about cost per accepted result, not the dramatic tokens-per-dollar headline.

Skip Groq as the only provider when the application requires a proprietary model absent from its catalog, a guaranteed best-effort route, or a preview model that has no stable replacement. Enterprise Performance can address contracted capacity for its supported models, but it is custom provisioned pricing rather than a cheaper public tier.

The explicit decision rule is this: choose Groq only when a current production model passes the acceptance test and either synchronous operating behavior or Batch cost beats the same-model fallback. If neither is true, the low token rate is not a reason to migrate.

Groq's price history rewards live verification

Groq pricing changes at the feature and model level, even when the billing structure stays pay as you go. The Developer launch in February 2025 advertised Batch at 25% off synchronous pricing. The current Batch documentation sets that discount at 50%.

Groq also lowered GPT-OSS prices and began adding prompt caching to the family in October 2025. The current result is $0.15/$0.60 for GPT-OSS 120B and $0.075/$0.30 for GPT-OSS 20B, with cached input at half the standard input rate.

Catalog churn matters just as much as price history. Models that appeared in older pricing coverage, including Kimi and DeepSeek routes, are absent from the current Supported Models page. A budget should carry its verification date, model ID, and fallback instead of treating a copied price as permanent.

The Monday move: price one accepted workflow and cap it

On Monday, take one bounded workflow and produce a budget that engineering and finance can both inspect. Do not begin with a provider-wide migration.

  1. Capture the traffic shape

    Export a representative week of request count, input tokens, output tokens, cache-hit tokens, model IDs, tool calls, retries, and accepted results. Separate synchronous traffic from jobs that can wait.

  2. Price three routes

    Calculate On Demand at the live model rates, Batch at 50% for eligible delayed work, and tools as their own line items. Do not apply cache savings to Batch, and do not assume a cache ratio that usage fields have not measured.

  3. Set the paid boundary

    If Free limits interrupt the route, move it to Developer. Set a monthly Spend Limit 20% to 30% above the forecast, then place alerts at 50%, 75%, and 90%. Leave room for the 10 to 15 minute reporting delay.

  4. Keep one fallback warm

    Run the same model through one alternate provider for a small evaluation set. Record the endpoint, observed acceptance rate, and reroute trigger so a catalog change or capacity error does not become an emergency migration.

The result is one budget line per accepted workflow: model tokens, tools, retries, and human review. Re-run it when the live model catalog, rate limits, or price page changes.

Frequently asked questions

How much does Groq AI cost?

Groq Free costs $0. Developer has no base subscription and charges usage by model. Current production text prices run from $0.05 input and $0.08 output per million tokens for Llama 3.1 8B Instant to $0.59 input and $0.79 output for Llama 3.3 70B Versatile.

Is Groq free to use?

Yes, within model-specific organization limits. GPT-OSS 120B and 20B currently list 30 RPM, 1,000 RPD, 8,000 TPM, and 200,000 TPD on Free. Exceeding an applicable limit returns HTTP 429.

Which Groq AI models are free to use?

The Free plan exposes the supported catalog under model-specific caps. Use production models for a durable evaluation. Groq says preview models are for evaluation and may be discontinued at short notice.

How much is Groq per month?

There is no public monthly or annual self-serve subscription. Add monthly input, output, audio, tool, and service-tier usage at the current rates. The worked 100,000-request GPT-OSS 120B workload costs $60 synchronously before cache or tool charges.

Does Groq offer a student discount?

Groq does not list a student-specific discount on the pricing or billing pages verified in August 2026. Students and other light users can use the general $0 Free plan within its rate limits.

What is Groq's refund policy?

Groq handles general billing refund requests case by case through support. Service Credits are stricter: purchased and promotional credits are non-refundable, non-transferable, and expire after one year unless a different term is stated.

Did Groq pricing change?

Yes. Batch launched at 25% off synchronous pricing in February 2025 and is 50% off today. Groq lowered GPT-OSS pricing and added caching to the family in October 2025. Verify the live page because the catalog and model-level rates move.

How does Groq compare with OpenAI pricing?

OpenAI does not serve GPT-OSS through its own API, so a direct comparison changes the model. For a normalized GPT-OSS 120B workload, Groq, Together AI, and Fireworks AI all cost $27 for 100 million input and 20 million output tokens before cache discounts.

Get the AI Tools Map for Business Owners

The AI Tools Map for Business Owners turns model prices into an adoption route, with the quality floor, operating role, and cost trigger each tool must earn. Subscribe to get the next edition free.

Last Updated

Aug 15, 2026

CategoryAI
Newsletter

One letter, every Sunday. Working systems, not hot takes.

Build logs, working systems, and field notes from running a portfolio of AI ventures.

Weekly. No spam. Unsubscribe anytime.