Cheapest AI API
Eight low-cost AI APIs compared with live August 2026 prices, free tiers, production limits, and the monthly cost that changes the decision.
Published Updated

DeepInfra has the lowest sticker price in this comparison: Mistral Nemo costs $0.019 per million input tokens and $0.03 per million output tokens. For a cheap production app, GPT-6 Luna and Claude Haiku 5.5 now tie at $0.10 input/$0.50 output on short prompts; keep the provider your app already uses. Gemini 2.5 Flash-Lite still wins for a free prototype, while Haiku loses its price tie above 100,000 prompt tokens; the cheapest AI API for your app is the lowest rung that clears its quality, privacy and latency floor.
Cheapest AI API: the verdict at a glance
GPT-6 Luna is the first cheap production candidate for an existing OpenAI app; Claude Haiku 5.5 is its price-matched alternative for a Claude app with prompts up to 100,000 tokens. Both cost $2 for 10 million input tokens plus 2 million output tokens on Standard, or $1 through Batch. Direct DeepSeek is no longer the automatic budget default: its current V4.1 Flash costs $2.70 off-peak or $5.40 at peak for that workload.
"Cheapest" still describes different jobs:
- DeepInfra is cheapest on raw token price. Mistral Nemo costs $0.019 per million input tokens and $0.03 per million output tokens.
- Google Gemini is cheapest for a prototype. Gemini 2.5 Flash-Lite has a free tier, then costs $0.10 input and $0.40 output per million tokens on Paid Standard.
- OpenAI and Anthropic are the first budget candidates inside an existing mainstream app. GPT-6 Luna and short-prompt Haiku 5.5 tie at $0.10/$0.50. The pricing tie is a reason to keep a working integration, not evidence that the models have identical quality.
- DeepSeek remains worth evaluating for long, repeated context. V4.1 Flash retains JSON, tools and a 1M-token context window, with cache-hit input at $0.003 off-peak or $0.006 at peak. Price the output and time of day too.
Since October 7, 2026, Claude Haiku 5.5 has brought Anthropic into the $0.10-input/$0.50-output price class. Anthropic's pricing page limits those rates to prompts up to 100,000 tokens; longer prompts cost $0.50/$2.50. The short-prompt price matches GPT-6 Luna's official API rate, so this launch changes which providers belong on the shortlist without making Haiku the cheapest route for every prompt length.
Every price below was verified against the provider's live pricing or documentation page on October 8, 2026. Prices are per 1 million input/output tokens unless the cell says otherwise. Stored pricing screenshots and illustrations are earlier snapshots; the refreshed tables and linked primary pages carry the current rates.
The table is not a quality ranking. Ministral 3 3B at $0.10/$0.10 and a current model at $0.10/$0.50 are different products even when both accept a chat-completions request. Use sticker price to choose candidates, then use your own acceptance test to choose the winner. Mistral's tiny-model entry rate is distinct from the Small 4 route priced in the workload table.
What 20,000 AI requests cost
At small and medium volumes, every provider here is cheap enough that one failed workflow decision matters more than the token bill. The useful comparison is still a normalized workload because it exposes output-heavy pricing and fees that a single input rate hides.
Assume a product makes 20,000 calls a month. Each call contains 500 billable input tokens and uses 100 billable output tokens, including any charged reasoning output. That is 10 million input tokens plus 2 million output tokens:
Monthly cost = 10 ร input price + 2 ร output price.
The table orders inference cost for these named routes. Standard rates apply unless a discount or time window is stated. It excludes cache savings, cache storage/writes, paid tools, regional premiums, retries and gateway funding fees. Equal token counts are a billing comparison, not a claim that every model tokenizes the same text identically.

The spread looks large in percentage terms and tiny in absolute dollars. OpenAI Luna costs $1.75 more than DeepInfra Mistral Nemo for this workload. Direct DeepSeek costs $0.70 more than Luna or short-prompt Haiku off-peak and $3.40 more at peak. If the cheaper model breaks a schema, misses a customer request, or creates a manual review queue, the saved dollars disappear quickly.
That is the quality floor: the minimum model behavior your task can accept without compensating work. A classification endpoint with five allowed labels has a low floor. A customer-facing support answer grounded in account data has a much higher one. An agent allowed to update records has a higher floor again because the cost of a bad action is not another API call.
Output pricing deserves special attention. A summarizer may read a long document and return a short result, so input dominates. A coding agent, sales-email generator, or conversational assistant can produce far more output. Luna and short-prompt Haiku both charge $0.50 for output, five times their $0.10 input rate; Mistral Small 4's $0.60 output rate is four times its $0.15 input rate. Flat input/output pricing on tiny models looks attractive in output-heavy work, but only if those models finish the task.
1. OpenAI: the best cheap default for an existing OpenAI app
OpenAI is now a first budget candidate for an app already built around its API. GPT-6 Luna costs $0.10 per million input tokens, $0.01 cached input and $0.50 output at the headline Standard rate. OpenAI's own API pricing page scopes that rate to context lengths under 272K and gives Batch a 50% input/output discount. The normalized workload is $2 Standard or $1 Batch.

Migration has a cost even when it never appears on an API invoice. Existing schemas, prompts, moderation rules, tools, tracing, fallback behavior and workload evaluations can make a small token difference trivial. With Haiku at the same short-prompt rate and direct DeepSeek costing more for uncached requests, the first move is to evaluate Luna inside the integration you already have.
Best for: Existing OpenAI products, cached stable prompts and asynchronous batch jobs
Standout: A $2 Standard workload that ties short-prompt Haiku, with a broader headline context-price range
Pricing: Standard $0.10 input/$0.01 cached/$0.50 output; Batch input/output $0.05/$0.25
Free trial: No standing free API tier listed on the pricing page
- $1 normalized cost through Batch
- Cached input costs one tenth of the headline Standard input rate
- Existing OpenAI applications can evaluate a cheaper model without a provider migration
- Published headline Standard pricing extends under 272K context, beyond Haiku's 100K low-price prompt tier
- Output costs five times input at the headline Standard rate
- The headline does not price every possible context length
- Optional data residency adds a 10% pricing premium
- Batch waits for asynchronous processing; it is not the interactive Standard route
Optimize inside OpenAI before migrating
The earlier OpenAI routing review explains the provider-ladder approach. For this cost decision, the first move is simpler: put short, repeatable work on GPT-6 Luna; cache eligible stable prefixes; send non-urgent work through Batch; preserve the stronger current model as fallback.
For 10 million input and 2 million output tokens, Luna Standard costs $2 and Batch costs $1. Short-prompt Haiku has the same two totals. Direct DeepSeek V4.1 Flash costs $2.70 off-peak or $5.40 at peak before cache savings. A working OpenAI application no longer needs to leave the provider to reach this price class.
Context and processing settings can change the result. OpenAI's advertised rate applies under 272K context, while Haiku's price increases above 100K prompt tokens. Do not extend either headline outside its stated scope. Flex is also listed as a lower-cost option with slower responses and occasional resource unavailability, but its exact Luna rate is not quoted in this comparison.
Pick OpenAI when the product already works there or Batch suits the queue. Skip Standard for offline jobs that qualify for half-price Batch, and reprice any request outside the headline context range.
2. Anthropic: the cheapest Claude API route for short prompts
Anthropic's Claude Haiku 5.5 is the budget starting point for an existing Claude workflow with prompts up to 100,000 tokens. Its $0.10 input/$0.50 output rate produces the same $2 normalized bill as GPT-6 Luna. Anthropic positions Haiku for high-volume classification, extraction and routing; those bounded tasks are the right place to evaluate it before moving customer-facing or action-taking work.
This price class changes the decision for a business that already has Claude prompts, validation and a stronger fallback in place. There is no token-price reason to migrate that short-prompt workload to Luna: Standard and Batch tie. There is also no reason to treat the launch as proof that Haiku passes every task your larger model handles.
Best for: Existing Claude apps, high-volume bounded text tasks and short-prompt batch queues
Standout: Anthropic at the same short-prompt input/output rates as GPT-6 Luna
Pricing: Standard $0.10 input/$0.50 output up to 100K prompt tokens; $0.50/$2.50 above 100K; short-prompt cache hits $0.01
Free trial: The API pricing docs list a small new-user testing credit without a fixed dollar amount
- $2 normalized Standard cost for the short-prompt workload
- Batch halves input/output prices to $0.05/$0.25 on short prompts
- Eligible short-prompt cache reads cost $0.01 per million input tokens
- A cheap route inside the same provider as an existing Claude fallback
- Prompts over 100K tokens raise both input and output rates fivefold
- Longer-prompt cache reads are $0.05, not the $0.01 headline
- Cache writes are charged separately; a read rate is not the first-call input rate
- Equal prices with Luna do not establish equal quality or identical token counts for the same text
Where the 100K tier changes the ranking
Haiku loses its price tie as soon as the prompt exceeds 100,000 tokens. Input moves from $0.10 to $0.50 per million and output from $0.50 to $2.50; Batch remains half-price, at $0.25/$1.25 for those longer prompts. These are prompt-length tiers, not a cheap input rate that can be applied to every document.
Take one uncached Standard request with 150,000 input tokens and 100 billed output tokens. Haiku costs $0.07525. Luna, still inside its advertised context-price range, costs $0.01505. Direct DeepSeek V4.1 Flash costs $0.02256 off-peak or $0.04512 at peak. Haiku ties Luna and undercuts uncached direct DeepSeek on the short workload, then becomes more expensive than both for this long prompt. The 20,000-request table uses 500-token inputs, so it never enters that tier.
Prompt caching, which reuses previously processed stable input, can reduce repeated-context cost. The pricing page lists Haiku cache reads at $0.01 per million up to 100K prompt tokens and $0.05 above. Five-minute cache writes cost $0.125 and $0.625 respectively. Count actual cache hits and writes rather than applying the read rate to fresh input.
Pick Anthropic when a Claude integration already works and the task stays in Haiku's low-price prompt tier. Skip the $0.10/$0.50 forecast for long documents or accumulated conversations above 100K tokens.
3. Google Gemini: the best free AI API
Google Gemini is the best free route for a prototype that uses sanitized data. Gemini 2.5 Flash-Lite provides free input and output tokens on the Free tier. When the app moves to Paid, Standard pricing is $0.10 per million text, image, or video input tokens and $0.40 per million output tokens; Batch and Flex halve those rates to $0.05/$0.20.

The free tier has a privacy consequence that belongs in the decision, not the footnotes. Google's pricing page says Free-tier content is used to improve its products. Paid-tier content is not. Free is a sensible route for synthetic prompts, public documents, and a developer's test set; it is not the default for customer records, confidential documents, or unreleased code.
Best for: Free prototypes, multimodal inputs, batch processing, and cost-sensitive document workflows
Standout: Free input/output route plus paid text, image, video, and audio input
Pricing: Free tier; Paid Standard $0.10 text/image/video input and $0.40 output; Batch/Flex $0.05/$0.20
Free trial: Yes, a Free tier with limited model access
- Free input and output tokens on the Free tier
- $1.80 for the normalized workload on Paid Standard
- $0.90 for the same normalized workload on Batch or Flex
- Paid tier includes context caching and does not use content to improve Google products
- Free-tier content is used to improve Google products
- Free access is limited by model and quota
- Audio input costs $0.30 per million tokens on Standard, three times text/image/video input
The free-to-paid handoff
A founder validating a document-tagging feature can use the Free tier with public or generated documents, verify the response shape, and learn the token profile before adding billing. Once customer documents enter the flow, moving to Paid is not merely about volume. It changes the content-use term and provides paid context caching and Batch access.
For the normalized workload, Paid Standard costs $1.80. Batch or Flex costs $0.90. That makes Gemini cheaper than DeepSeek for asynchronous work on sticker price, although the correct choice still depends on which model clears the task. If results do not need to return immediately, the 50% reduction should be evaluated before changing provider.
Gemini also accepts image and video input at the same $0.10 Standard input rate as text for 2.5 Flash-Lite. That makes it a stronger candidate than text-only small models when a workflow reads receipts, screenshots, or product imagery. The pricing page still lists 2.5 Flash-Lite as a small, cost-effective model for at-scale use. It remains a cheap starting tier rather than a claim that this older version matches every newer Gemini model.
Pick Gemini when the project needs a free start, multimodal input, or half-price asynchronous processing. Skip the Free tier for sensitive data, and skip Flash-Lite when the acceptance test shows that a larger model is required.
4. DeepInfra: the cheapest raw token price
DeepInfra is the raw-price winner at $0.019 input and $0.03 output per million tokens for Mistral Nemo Instruct. The normalized 20,000-request workload costs $0.25. DeepInfra also lists Meta Llama 3.1 8B at $0.02/$0.04 and DeepSeek V4 Flash at $0.09/$0.18, so a builder can climb the same host's model ladder without opening another billing account.

The wall is the model behind the price. Mistral Nemo and Llama 3.1 8B are budget open-model routes; the price alone does not establish their acceptance rate on your task. A rigid labeling task may need nothing more. A customer-facing workflow should not inherit their low price until its own evaluation shows that the output is acceptable.
Best for: High-volume low-risk classification, tagging, lightweight extraction, and model price experiments
Standout: The lowest published text-generation token rate in this comparison
Pricing: Mistral Nemo $0.019 input/$0.03 output; Llama 3.1 8B $0.02/$0.04
Free trial: No; DeepInfra says a card or prepayment is required
- $0.25 for the normalized workload on Mistral Nemo
- Large catalog spanning text, embeddings, image, audio, and video models
- Standard, Priority, and lower-cost Flex scheduling
- Automatic scaling with a published 200-concurrent-request account limit
- The lowest rate buys a small model, not a universal production default
- Card or prepayment is required
- Flex can be slower and occasionally unavailable
Use DeepInfra when the answer space is narrow
A local-services marketplace that labels inbound messages as quote, reschedule, cancel, billing, or other has a narrow output space. It can reject anything outside five values and send ambiguous cases to a fallback. That is the kind of job where a $0.019/$0.03 model deserves an evaluation.
A support assistant answering account-specific questions is different. It must retrieve the right record, follow policy, preserve nuance, and avoid a confident wrong answer. Saving $1.75 against Luna or short-prompt Haiku on the normalized workload fails to justify moving that job to Mistral Nemo without a stronger acceptance result.
DeepInfra's service tiers create another lever. Standard is the 1ร base rate. Priority is 1.5ร for faster scheduling during peak demand. Flex is 0.8ร in exchange for slower responses and occasional unavailability. Where the selected model supports it, use Flex for backfills, offline enrichment, or a retryable nightly queue. Do not put an interactive customer response on it merely to shave 20% off a bill that is already measured in cents.
Pick DeepInfra when a narrow task and a strong validator make small-model savings repeatable. Skip the cheapest row for open-ended writing, complex reasoning, and actions that can change customer or business data.
5. Groq: cheap when latency matters
Groq is the low-cost pick to evaluate when an interactive response must feel immediate. Its production GPT OSS 20B route costs $0.075 per million input tokens and $0.30 per million output tokens, with a published speed of about 1,000 tokens per second. The normalized workload costs $1.35, or $1.10 more than DeepInfra's raw-price winner. Published generation speed is not a measurement of your end-to-end latency.

The old Llama 3.1 8B Instant price cannot support the comparison anymore. Groq's current production table marks that route Enterprise and Contact Sales. Its old numeric row is removed; GPT OSS 20B is the publicly priced production alternative used here. Preview models remain a separate evaluation lane and may be discontinued at short notice.
Best for: Latency-sensitive classification, autocomplete and fast first-pass routing
Standout: Published speed of about 1,000 output tokens per second on GPT OSS 20B
Pricing: Production GPT OSS 20B $0.075 input/$0.30 output per million tokens
Free trial: Public trial amounts are not specified on the documentation pages checked
- $1.35 for the normalized GPT OSS 20B workload
- About 1,000 tokens per second listed by the provider
- Mostly OpenAI-compatible client libraries reduce the first integration step
- Public documentation distinguishes production and preview models
- Llama 8B no longer has a public numeric rate on the production table
- Preview models may disappear at short notice
- API compatibility does not mean identical model behavior or support for every OpenAI field
- Model speed does not include your network, queue and application overhead
Why latency can be the cheaper system
A customer-service form that drafts a short answer or a sales inbox that labels a message benefits when the response appears immediately. The bill difference between a fast route and the absolute token floor is small at this volume. A route that keeps the interface usable can be worth more than the final dollar saved.
GPT OSS 20B costs less on this workload than Luna or short-prompt Haiku, but the choice still belongs to the same representative evaluation set. Lower cost and higher published token speed do not establish better instructions, tools or business fit.
Pick Groq when response time is visible to the user and its publicly priced production model passes the task. Skip a sales-priced or preview route when your budget requires a reproducible public unit rate.
6. Mistral: the cheapest direct small-model lab
Mistral is the cleanest direct-vendor ladder for small models. Ministral 3 3B starts at $0.10 input and $0.10 output per million tokens, the 8B model is $0.15/$0.15, and the 14B model is $0.20/$0.20. Mistral Small 4 costs $0.15/$0.60 and adds function calling, structured outputs and agents to the product surface.

The cheap 3B row is not the default recommendation for open-ended customer work. Use it where a tiny model has an advantage, such as classification, moderation pre-routing, or a constrained field transform. Mistral Small 4 is the more credible general application candidate and costs $2.70 for the normalized workload.
Best for: Direct access to small models, direct-vendor preference, batch transforms, and cached repeated prompts
Standout: Flat $0.10/$0.10 entry price plus 50% Batch and 90% cached-input discounts
Pricing: Ministral 3 3B $0.10/$0.10; Mistral Small 4 $0.15/$0.60
Free trial: No broad free text-generation tier is listed; Mistral Moderation 2 is a separate free endpoint
- Simple direct-vendor pricing across 3B, 8B, 14B, and Small tiers
- 50% batch discount for high-volume asynchronous work
- 90% cached-input discount for repeated stable prefixes
- Mistral Small 4 lists function calling, structured outputs and batching
- The $0.10/$0.10 headline belongs to a 3B model
- Mistral Small 4 output is four times its input price
- No broad free production tier is listed for the main text models
When the Mistral ladder pays
Suppose an ecommerce platform generates nightly catalog tags from titles, attributes, and descriptions. The result is asynchronous, each item follows a stable schema, and a human can inspect sampled errors before the next catalog push. That workflow can begin on Ministral 3 3B or 8B, use Batch for 50% off, and evaluate eligible cache reads for the stable taxonomy instructions.
The economics change for a chat assistant. On the normalized workload, Ministral 3 3B costs $1.20, while Mistral Small 4 costs $2.70. The $1.50 difference buys the broader Small 4 API surface. If Small 4 reduces fallback and review, it can be the cheaper system even though its token line is higher.
Pick Mistral when a direct model lab or small-model ladder matters. Skip the 3B headline rate when the task is open-ended and the evaluation has not proved it can stay above the quality floor.
7. SiliconFlow: the low-cost multi-model newcomer
SiliconFlow is a low-cost multi-model host with a concrete starting-credit budget. Its GPT OSS 20B price is $0.04 per million input tokens and $0.18 per million output tokens with a 131K context window. That produces a $0.76 bill for the normalized workload, and new accounts receive $1 in credits.

The important distinction is the model name. GPT OSS 20B is the hosted open-model route compared here; it is separate from the direct GPT-6 Luna API. SiliconFlow also lists DeepSeek V4.1 Flash at $0.15 input, $0.003 cached input, and $0.60 output, so the catalog can provide a stronger step-up route without abandoning the account.
Best for: Cost-sensitive builders who want several open models behind one provider account
Standout: GPT OSS 20B at $0.04/$0.18 with $1 in starting credits
Pricing: GPT OSS 20B $0.04 input/$0.18 output; DeepSeek V4.1 Flash $0.15/$0.60
Free trial: $1 in free credits and no minimum commitment
- $0.76 for the normalized GPT OSS 20B workload
- $1 in credits covers more than that normalized test workload at list price
- No minimum commitment on the public pricing page
- Low-cost routes across DeepSeek, Qwen, MiniMax, GPT OSS, and other open models
- The lowest prices belong to open-weight models, not proprietary frontier APIs
- Model choice still transfers quality evaluation to the buyer
- A multi-model host may need extra procurement and reliability review
The $1 evaluation budget
SiliconFlow's starting credit is unusually concrete. The normalized GPT OSS 20B workload costs $0.76, so the $1 credit can cover one full price-normalized run before retries or other models. That does not mean a free production month. It means a builder can measure output shape, latency, and token accounting without precommitting to a larger balance.
For a back-office summarization queue, start with GPT OSS 20B and send failed or low-confidence cases to DeepSeek V4.1 Flash on the same platform. The model rates make that ladder legible: $0.04/$0.18 for the cheaper route, then $0.15/$0.60 for DeepSeek. The application still needs to log which model answered because a blended bill without model-level acceptance data cannot reveal whether the cheaper tier is helping.
Pick SiliconFlow when one low-cost host and a broad open-model catalog simplify evaluation. Skip it when a buyer requires a direct contract with the model developer or already has a gateway that provides the same models at an acceptable total cost.
8. OpenRouter: the cheapest way to keep switching
OpenRouter is the best budget gateway when model switching is worth a small billing fee. It passes through provider inference prices without a markup, exposes free model variants, and places many providers behind one API key. The current DeepSeek V4.1 Flash provider table offers Baidu Qianfan at a promotional $0.0354 input/$0.1416 output per million tokens. That specific pair costs $0.6372 for the normalized workload before Standard-plan credit-purchase fees.

The wall is the way credits are funded. OpenRouter Standard charges 5.5% with a $0.80 minimum when credits are purchased. Its current plan pricing distinguishes Standard from Business, which lists an 8% platform fee. The minimum stops binding at about $14.55, because $0.80 divided by 5.5% is $14.5455. A $10 credit purchase therefore costs a $0.80 fee, or 8%, rather than 5.5%.
Best for: Comparing models, keeping fallback routes, and changing providers without rewriting the application
Standout: Provider prices passed through without inference markup
Pricing: Free variants at $0; paid models vary; DeepSeek V4.1 Flash via Baidu Qianfan promotional $0.0354/$0.1416 before the Standard credit fee
Free trial: Small new-user allowance plus free variants with low daily limits
- One API key across many providers and models
- No markup on the underlying inference rate
- Free variants for evaluation
- Standard includes $25,000 of list-price inference each month without a BYOK fee
- Standard credit purchases cost 5.5% with a $0.80 minimum
- Free variants allow only 50 requests per day before a $10 lifetime credit purchase, then 1,000 per day
- OpenRouter reserves the right to expire unused credits after one year
- Gateway and upstream-provider failures both belong in the reliability plan
When the gateway fee pays back
The model's provider pricing table matters more than the catalog's lowest input number. Relace lists $0.0283 input but $1 output, producing $2.283 for this workload. Baidu Qianfan's current promotion has a higher input price but much cheaper output, producing $0.6372. Pin that provider to reproduce the promotional rate; automatic fallback can change the price. No promotion end date is stated, so do not use it as a fixed annual assumption.
A builder trying DeepSeek, Gemini, and an open model can spend more engineering time on separate auth, billing, retry behavior, and response formats than on inference. OpenRouter makes that switch a routing configuration. The fee is reasonable when that convenience keeps a fallback ready or lets a team replace a weak route quickly.
Tiny Standard-plan top-ups distort the economics. At $10, the $0.80 minimum is an 8% fee. At $100, 5.5% is $5.50 and the percentage behaves as advertised. A small prototype should either accept the minimum as a convenience charge or use a provider's direct free tier. Do not compare OpenRouter's model rate to a direct rate and forget the credit fee.
The BYOK route has a different rule. Standard includes $25,000 of list-price inference per month using your own provider keys without a BYOK fee, then charges 5% above that allowance. The current allowance is measured in inference cost, not request count. BYOK can preserve direct provider limits while centralizing routing, but it does not remove the need to monitor two billing and failure surfaces.
Pick OpenRouter when one integration and rapid switching are worth the fee. Skip it when the product uses one stable provider, the direct API already meets uptime needs, and the extra gateway adds no operational value.
9. DeepSeek: long context and cache-heavy work, with time-based pricing
DeepSeek remains a candidate when an app needs JSON, tools and long context, but it is no longer the cheapest uncached direct default. The current API model is V4.1 Flash, called as deepseek-flash. It costs $0.15 input/$0.60 output off-peak and $0.30/$1.20 at peak. Its pricing page lists a 1M-token context, up to 384K output, JSON output, tool calls, Responses API, an Anthropic-format route and vision.

The warning about a future price increase has been overtaken by the August 16, 2026 pricing change and the September 10 V4.1 Flash replacement. The current pricing page says the old V4 Flash aliases are served by V4.1 Flash and billed at its rates. Keeping an old model string does not keep the old introductory price.
Best for: Production extraction, structured generation, economical agents, and long-context jobs
Standout: 1M context plus JSON, tools, vision, Responses API and Anthropic-format API support
Pricing: Off-peak $0.15 input/$0.003 cache-hit/$0.60 output; peak $0.30/$0.006/$1.20 per 1M tokens
Free trial: No standing free tier listed on the pricing page
- $2.70 for the normalized workload when every request is off-peak
- 1M-token context and up to 384K output
- JSON output and tool calls at the budget tier
- Published concurrency limit of 2,500 for V4.1 Flash
- Peak pricing doubles the uncached normalized bill to $5.40
- A standing free tier is not published
- A low-cost external route still needs privacy, reliability, and regional review for the buyer's use case
Where DeepSeek remains worth evaluating
DeepSeek remains a candidate when the task needs structured output over long or repeated context. Consider a B2B SaaS product that extracts renewal date, contract value, notice period, and named risks from uploaded agreements. The output can be checked against a schema, but the model still needs enough context to read a long document and enough instruction following to return all required fields. V4.1 Flash offers an API surface worth including in the evaluation, especially when prompts cross Haiku's cheap-context tier.
Cache-hit input costs $0.003 off-peak or $0.006 at peak. That matters when the same eligible instructions or reference prefix recur. It does not make fresh input cost that rate, and a high output bill can still outweigh the cache saving. Budget cache misses at the relevant $0.15 or $0.30 rate until usage data proves the hit ratio.
Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday through Friday, excluding Chinese public holidays. All other hours are off-peak, including whole weekends and Chinese public holidays. A queue that can run outside those windows costs $2.70 rather than $5.40 for the uncached normalized workload. Live traffic must be budgeted for the hours in which it actually arrives.
A one-week DeepSeek evaluation
Choose one bounded task
Start with extraction, classification, summarization, or a read-only agent action. Define the fields, allowed outputs, and failure conditions before sending traffic.
Build a 100-request evaluation set
Use representative, sanitized requests from the workload. Include short, long, ambiguous, malformed, and edge-case inputs so the cheapest route cannot win by seeing only easy examples.
Require machine-checkable output
Use JSON output where the job has a schema. Record valid-schema rate, accepted-answer rate, latency, input tokens, output tokens, and cache hits for every call.
Keep the current model as fallback
Send failed validation, low-confidence cases, and any action with material consequence to the existing trusted route. The first rollout should be reversible.
Record the pricing window
Log when each request runs, its cache hits and the charged model version. Recalculate from current peak/off-peak rates when the vendor changes prices or redirects a model alias.
Pick DeepSeek if V4.1 Flash passes a long-context or repeated-context task and its full bill beats the alternatives. Skip it when a fixed annual unit price, a specific processing region, or an established vendor relationship outweighs the token difference.
Who should pick what
Choose the route by consequence first, then price the cheapest model that survives it. That one rule prevents most false savings.

For a free sanitized prototype, pick Gemini 2.5 Flash-Lite. It gives the cleanest zero-dollar start and a straightforward move to $0.10/$0.40 Paid Standard. Keep customer and confidential data out of Free because its content-use term differs from Paid.
For bounded live production in an existing mainstream app, start with Luna or short-prompt Haiku. Both cost $2 for the normalized Standard workload. Keep OpenAI if it already works there; keep Claude if Haiku passes and prompts stay at or below 100K tokens. For longer prompts, Haiku loses that price tie.
For a narrow high-volume labeler, evaluate DeepInfra first. The $0.25 normalized cost is meaningful only where allowed outputs are constrained and failures are cheap to detect. Move up immediately when the task becomes open-ended.
For visible latency, evaluate Groq GPT OSS 20B. Groq lists about 1,000 tokens per second and $0.075/$0.30 on its production route. Use it as a fast lane with a deeper fallback. Llama 8B now requires a sales quote, so its old rate cannot decide the purchase.
For offline queues, compare Gemini Batch/Flex, Mistral Batch, OpenAI Batch and Anthropic Batch. Each quoted discounted route documents a 50% reduction. Price Haiku's long-prompt Batch tier separately, and schedule DeepSeek off-peak when using it. The best provider is the one whose model passes the queue's acceptance check, not the one with the lowest undiscounted Standard line.
For easy switching, pick OpenRouter. Its fee can be cheaper than maintaining multiple integrations, especially when fallback and model churn are part of the product. For one stable provider, direct billing is simpler.
For an existing OpenAI application, try GPT-6 Luna before migrating. Batch brings the normalized cost to $1, matching short-prompt Haiku Batch and below uncached direct DeepSeek even off-peak. Existing evaluations and operations strengthen the case for staying.
For long, repeated context, evaluate DeepSeek V4.1 Flash. Its 1M context and low cache-hit input rates remain useful, but fresh input and output cost $0.15/$0.60 off-peak or $0.30/$1.20 at peak. Measure the complete bill using the current model and rates.
The decision flips when the cheaper route falls below the workload's acceptance threshold. Keep the lower-priced model while schema validity, accepted answers, latency, privacy terms, and recovery behavior meet the requirement. Move up a rung when one of those breaks. Do not keep adding prompt patches to save cents.
A builder exposing an API-driven experience to customers also has to own auth, billing, review, and failure handling. The v0 API build guide shows why the underlying model or agent is only one layer of the product.
How these APIs were picked
The ranking rewards the cheapest dependable buying decision, not the lowest isolated input number. Each provider was priced and analyzed against its live official pages on October 8, 2026. The APIs were not subscribed to or presented as if they had been exercised in production.
The decision uses these criteria:
- Useful entry price: input and output rates for a named, currently listed model
- Quality floor: whether the model tier is plausible for a constrained task or a broader production workflow
- Operational surface: JSON, tools, context, batch, caching, rate limits, and fallback implications
- Privacy and contract consequence: how free data is handled, whether a price is stable, and whether a separate gateway fee applies
- Switching cost: the work required to migrate prompts, schemas, evaluation sets, monitoring, and billing
Nine providers made the cut because each owns a distinct buyer decision. Numbered picks are buying recommendations; the cost table separately ranks the selected inference routes. Groq's unpriced Llama 8B row was cut and replaced with its publicly priced GPT OSS 20B route. Old DeepSeek and OpenRouter V4 routes and the old OpenAI Luna version were replaced with models and rates verified today. A provider was excluded when its cheapest production text route sat well above this cost band, its price could not be normalized per token, or its role duplicated another option without a clearer benefit. Specialized image, video, speech, and embedding APIs were also excluded because their units are not comparable to text input/output tokens.
The cost math is original analysis using one explicit workload. It is not a benchmark. Your accepted-output cost can only come from your own representative requests, validators, and review process.
The ones to avoid
OpenRouter free variants as a production dependency
Avoid making an OpenRouter free variant the only route behind live customer traffic. The account gets 50 free-model requests per day until it has purchased at least $10 in credits, then 1,000 per day. That is an evaluation allowance, not a dependable capacity plan.
Gemini Free with customer or confidential data
Avoid sending sensitive content through Gemini Free. Google says Free-tier content is used to improve its products, while Paid-tier content is not. The normalized Paid Standard bill is $1.80, so the privacy boundary is not a sensible place to save that amount.
DeepInfra Mistral Nemo for unbounded customer-facing work
Avoid promoting the $0.019/$0.03 route straight into open-ended support, legal interpretation, or autonomous actions. The rate belongs to a small model. Use it for a constrained, validated job and keep a stronger fallback until the acceptance data earns broader scope.
Haiku above 100K priced as a short prompt
Avoid applying Haiku's $0.10/$0.50 rate to prompts over 100K tokens. Those requests use $0.50/$2.50 Standard, and their cache-hit input is $0.05 rather than $0.01. The threshold can make Luna or direct DeepSeek cheaper for the same long prompt.
DeepSeek without a pricing-window and model-version check
Avoid forecasting the retired V4 introductory rate or treating all hours as off-peak. Today's V4.1 Flash is $0.15/$0.60 off-peak and $0.30/$1.20 at peak, including requests through the old V4 Flash aliases. Log the actual pricing window and retain a fallback.
OpenAI Standard for retryable offline jobs
Avoid paying Luna's $0.10/$0.50 Standard rate for work that qualifies for Batch. Its 50% input/output reduction cuts the normalized workload from $2 to $1 without changing provider.
The Monday move
Next week, move one bounded task class, not the whole application. Start with a 100-request evaluation to look for obvious schema, latency and routing failures before production traffic is involved, then continue monitoring beyond that sample.
Pull 100 representative requests
Choose a single task and sample normal, difficult, ambiguous, and malformed inputs. Remove personal, confidential, and customer data before using any free tier.
Run three routes
Compare the current route, GPT-6 Luna or short-prompt Haiku 5.5, and Gemini 2.5 Flash-Lite. Include DeepSeek V4.1 Flash when long context or cache-heavy traffic is the objective. Add DeepInfra or Groq only when raw price or latency is the explicit objective.
Score four outcomes
Record valid-schema rate, accepted-answer rate, end-to-end latency, and exact input/output token cost. Keep human review time as a separate operational measure.
Route one winning class
Send only the bounded task that passed to the cheaper provider. Preserve the old route for failed validation, ambiguous inputs, and higher-consequence actions.
Reprice on every vendor change
Save the formula and workload assumptions. Recalculate when prompt lengths cross Haiku's 100K tier, DeepSeek pricing windows or model aliases change, an OpenRouter promotion ends, or free-tier and gateway limits move.
That move turns a list of rates into a reversible budget decision. It also produces the one number a public price table cannot: your cost per accepted result.
Frequently asked questions
Which AI has the cheapest API?
DeepInfra has the lowest raw text-generation price in this comparison: Mistral Nemo at $0.019 per million input tokens and $0.03 per million output tokens. For an existing mainstream production app, evaluate GPT-6 Luna or short-prompt Claude Haiku 5.5 at $0.10/$0.50, then choose by acceptance and integration cost. Haiku costs more above 100K prompt tokens.
Is there any free API for AI?
Yes. Google Gemini provides free input and output tokens on its Free tier, and OpenRouter offers free model variants. Gemini Free may use content to improve Google products, while OpenRouter free variants are limited to 50 requests per day or 1,000 per day after at least $10 in lifetime credit purchases.
How much does an AI API cost?
For 20,000 monthly calls totaling 10M input and 2M output tokens, the selected inference routes run from $0.25 on DeepInfra Mistral Nemo to $5.40 on direct DeepSeek V4.1 Flash at peak. Luna and short-prompt Haiku each cost $2 Standard or $1 Batch; gateway funding fees are extra. Batch, Flex, caching, gateway fees, retries, and the model that clears the quality floor change the final bill.
How much does the ChatGPT API cost?
ChatGPT subscriptions and OpenAI API billing are separate. The budget OpenAI model compared here is GPT-6 Luna at $0.10 input and $0.50 output per million tokens at the headline Standard rate under 272K context, or $0.05/$0.25 input/output through Batch.
Download the AI business workflow audit checklist to turn this cost ladder into a one-week provider evaluation.
- Last Updated
- Category
- Build
- Language







