DeepSeek Pricing (2026): Free Chat, API Under $1/M Tokens

DeepSeek chat is free. See live V4 Flash and V4 Pro API rates, cache rules, refunds, break-even math, and 2026 alternatives.

Published Updated

Tools
  • DDeepSeek
DeepSeek Pricing (2026): Free Chat, API Under $1/M Tokens

DeepSeek chat costs $0. Its API now charges peak and off-peak rates: deepseek-flash costs $0.15 off-peak or $0.30 peak per million cache-miss input tokens, plus $0.60 or $1.20 for output; deepseek-v4-pro costs $0.66 or $1.32 for input, plus $1.98 or $3.96 for output. Default production work to Flash, schedule flexible work off-peak, and pay for Pro only when better accepted results justify its higher rate.

DeepSeek pricing at a glance

DeepSeek has one free consumer product and two paid API models. There is no monthly or annual self-serve API subscription: the API deducts token charges from your account balance as you use it.

These numbers were verified against DeepSeek's live Models & Pricing page on October 4, 2026.

Since August 16, 2026, at 16:00 UTC, DeepSeek has used peak and off-peak API pricing. Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday, excluding Chinese public holidays; every other hour is off-peak, including weekends and Chinese public holidays in full. The rates below are today's schedule, following the September 10 Flash update, rather than the launch-day prices.

Archived August 2026 capture of DeepSeek's API pricing page
August 2026 pricing-page capture; use the current rates in the table below
OptionInput per 1M tokensOutput per 1M tokensDecision
DeepSeek chatNot meteredNot meteredStay free for personal chat
deepseek-flash, off-peak$0.003 cache hit / $0.15 cache miss$0.60Default production model
deepseek-flash, peak$0.006 cache hit / $0.30 cache miss$1.20Same model, double the rate
deepseek-v4-pro, off-peak$0.022 cache hit / $0.66 cache miss$1.98Escalate the hard tail
deepseek-v4-pro, peak$0.044 cache hit / $1.32 cache miss$3.96Same model, double the rate

Move flexible batch work off-peak to halve the API token bill for the same model, token volume and cache-hit mix.

A token is the unit the model bills, not a request or a word. Input is everything sent to the model, including instructions, history, documents, and tool context. Output is everything the model returns. A cache hit is repeated input that DeepSeek can reuse at the lower rate; a cache miss is input it must process at the full rate.

The two API models share a 1M-token context window and a 384K maximum output. The pricing page lists no separate long-context surcharge inside those limits. Their operating limits differ: Flash lists 2,500 concurrent requests, while Pro lists 500, with limits applied per account.

DeepSeek chat is free, but the API is a separate bill

DeepSeek's consumer app costs nothing. Its official app announcement, read on October 4, 2026, still says it is 100% free, with no ads and no in-app purchases.

Free does not mean the API is free. The consumer app is an interface for a person; the API is the metered product software calls. DeepSeek's API page deducts usage from a granted or topped-up balance. A free chat account therefore does not turn programmatic traffic into free traffic.

The same product split catches buyers elsewhere: an app subscription and API billing can share a brand while remaining separate products, as the Perplexity pricing breakdown shows. For DeepSeek, the distinction is even cleaner because the app has no published subscription price while every API token has a line item.

Who never needs to pay

You never need to pay if your use is interactive, occasional, and satisfied by the official web or mobile app. That covers personal questions, ad hoc writing, document discussion, and students who do not need to connect DeepSeek to another product.

DeepSeek does not list a student-specific plan or discount on the pricing or app pages verified for this review. Ordinary consumer access already costs $0.

The boundary is automation. The moment a CRM, support workflow, internal tool, or customer-facing product needs to call DeepSeek without a person opening the app, budget for API tokens. Also avoid treating free chat as production infrastructure: the free-app page does not promise unlimited capacity or publish a service level.

Flash is the production default; V4 Pro is an escalation lane

DeepSeek V4.1 Flash should receive the first request in most production routes. It has the lower price and the higher listed concurrency limit. Both models now support the Responses API; Flash also supports vision, which Pro does not.

V4.1 Flash

Use deepseek-flash, which currently serves DeepSeek-V4.1-Flash. Off-peak, Flash charges $0.15 per million cache-miss input tokens, $0.003 for cache-hit input, and $0.60 for output. At peak hours those rates are $0.30, $0.006, and $1.20. It supports thinking and non-thinking modes, JSON output, tool calls, DeepSeek's Anthropic-compatible API, and the Responses API. The live page lists a 1M context window, 384K maximum output, and 2,500 concurrent requests.

The old deepseek-v4-flash and deepseek-v4-flash-vision-exp names remain accepted aliases, but their retired models have been replaced by V4.1 Flash and their requests are billed at the current Flash rate.

That combination makes Flash the sensible route for classification, extraction, summarization, first-draft generation, routine code changes, and bounded agent steps. These jobs have clear acceptance checks, so a failed output can be detected and escalated without letting a cheap mistake reach a customer.

Skip Flash as the only route when failure is expensive and hard to detect. A legal, financial, or customer-impacting recommendation should not choose its model from token price alone. The low rate buys room to evaluate; it does not replace evaluation.

V4 Pro

Use deepseek-v4-pro, which currently serves DeepSeek-V4-Pro-0813. Off-peak, Pro charges $0.66 per million cache-miss input tokens, $0.022 for cache-hit input, and $1.98 for output. At peak hours those rates are $1.32, $0.044, and $3.96. It has the same 1M context window and 384K maximum output, plus the same thinking modes, JSON output, tool calls, and Anthropic-compatible format. Its listed concurrency limit is 500.

Pro is the hard-tail route: the ambiguous request, the complex plan, the difficult coding task, or the second attempt after Flash fails a deterministic check. It is not the default merely because its name says Pro.

The current pricing table lists Responses API support for Pro as well as Flash. That removes the old endpoint limitation. Pro still lacks Flash's vision support, so an image-dependent workflow cannot treat the two as interchangeable.

The upside
What it does well
5 points

  • The consumer app costs $0, with no ads or in-app purchases.
  • Flash's off-peak cache-miss input and output rates stay below $1 per million tokens.
  • Both API models carry a 1M context window and 384K maximum output.
  • Automatic context caching can make repeated prefixes dramatically cheaper.
  • Flash combines the lowest rate, vision support, and higher concurrency.
The downside
Where it falls short
5 points

  • Peak rates double every API token line item.
  • Cache hits are best effort, not guaranteed.
  • Pro costs 4.4x Flash on cache-miss input and 3.3x on output at the same time band.
  • Pro has one-fifth of Flash's listed concurrency and no vision support.
  • Free chat has no published unlimited-capacity promise or service level.

DeepSeek's token bill is tiny; cost per accepted result still decides

The raw API expense is small enough that one avoidable human review can outweigh thousands of model calls. The useful formula is not cost per request. It is cost per accepted result.

For a request with no cache hits, use the rates for its pricing period:

cost = input tokens / 1,000,000 × input rate + output tokens / 1,000,000 × output rate

At 1,000 tokens, Flash costs $0.00015 for cache-miss input and $0.00060 for output off-peak, or $0.00030 and $0.00120 at peak. Pro costs $0.00066 for cache-miss input and $0.00198 for output off-peak, or $0.00132 and $0.00396 at peak.

Consider a support-drafting workflow with 2,000 input tokens and 500 output tokens per reply. At 1,000 replies, Flash costs $0.60 off-peak or $1.20 peak, and Pro costs $2.31 or $4.62, with no cache hits. The Pro premium is $1.71 off-peak or $3.42 peak per thousand replies.

Now add labor. Five minutes of review at an illustrative $60 hourly cost is $5. One extra review event costs more than the Pro premium for that entire batch in either period. The right question is therefore not, "Which model has the lower line item?" It is, "Which route clears the acceptance check with the least model spend plus review time?"

What a monthly DeepSeek bill looks like

DeepSeek has no seat-month to compare with a software subscription. Monthly cost is the traffic shape: how much context goes in, how much text comes out, how often prefixes hit cache, which hours the work runs, and how many attempts an accepted result takes. The examples below assume all usage falls in one time band; a mixed month needs each band's usage priced separately.

For a solo product prototype using 10M cache-miss input tokens and 2M output tokens in a month, Flash costs $2.70 off-peak or $5.40 peak, and Pro costs $10.56 or $21.12. The difference is $7.86 or $15.72. The model that produces a valid result with less debugging can still matter more than that gap.

For a support operation drafting 100,000 replies at 2,000 input and 500 output tokens each, the uncached bill is $60.00 off-peak or $120.00 peak on Flash, versus $231.00 or $462.00 on Pro. At an illustrative 80% input cache-hit rate, those totals fall to $36.48 or $72.96 on Flash and $128.92 or $257.84 on Pro. At off-peak rates, one hour of human review still exceeds the cached monthly Flash bill under the $60-per-hour assumption.

For 1,000 document-analysis jobs using 50,000 input and 2,000 output tokens each, the uncached bill is $8.70 off-peak or $17.40 peak on Flash, versus $36.96 or $73.92 on Pro. With 80% of input hitting cache, it becomes $2.82 or $5.64 on Flash and $11.44 or $22.88 on Pro. Long input does not automatically mean a large DeepSeek bill; repeated long input can make the unit economics unusually small.

These examples also show why a fixed monthly estimate copied from another company is useless. A support reply is output-heavy relative to its prompt. A document job is input-heavy. An agent can multiply both through retries and tool loops. Save the token mix and pricing period beside the dollar figure so finance can tell whether traffic, quality, scheduling, or a vendor rate caused the change.

The Flash-to-Pro break-even

Pro's uncached input rate is 4.4 times Flash, while its output rate is 3.3 times Flash. There is no single price multiple for every workload. With no cache hits and the same pricing period, the blended premium falls between those two ratios.

For the support example's 2,000-input and 500-output token mix, Pro costs 3.85x Flash in either period. Under that fixed mix, Pro must reduce billed attempts or token volume by 74.03% to break even on model spend alone. Change the input/output mix, cache-hit rate, or execution period and recalculate.

If Pro finishes a task in one attempt, Flash can average up to 3.85 comparable attempts before its token bill becomes higher in this example. Three Flash attempts cost less than one Pro attempt; four cost more. Human review can move that boundary sooner.

Historical August routing illustration with the previous Flash and Pro costs
August 2026 illustration: its prices and 3.11x label are historical; the current support-workload threshold is 3.85x

This routing rule also prevents a common procurement mistake. A team sees low input rates, sends every task to Pro, and calls the total cheap. The absolute bill may still be small, but an unmeasured premium across millions of calls is not a strategy. Flash should own the bounded base load; Pro should own the failures that justify it.

Context caching changes the bill only when prefixes truly repeat

DeepSeek's cache-hit rate is valuable, but it is not a discount you can assume in a spreadsheet. The context-caching guide says caching is enabled by default, works on complete persisted prefix units, and operates on a best-effort basis.

A stable system prompt followed by a stable reference document is a strong cache candidate. A later request can reuse that beginning and append a new question. Reusing a phrase from the middle of a different prompt does not qualify. The complete prefix must match a unit DeepSeek has persisted.

The cache also has a time dimension. Construction takes seconds, and unused entries are normally cleared within hours to days. A weekly report may miss a cache that a burst of near-identical requests would hit. DeepSeek does not guarantee a 100% hit rate.

At an illustrative 80% input cache-hit rate, the blended input cost per million tokens falls to $0.0324 off-peak or $0.0648 peak on Flash, and $0.1496 or $0.2992 on Pro. In the 1,000-reply support example, total cost falls from $0.60 to $0.3648 off-peak on Flash and from $2.31 to $1.2892 on Pro. Peak totals are $0.7296 and $2.5784 respectively.

That saves 39.20% of the total Flash bill and 44.19% of the Pro bill in this output-bearing workload. Output tokens do not receive the input cache discount, so an 80% input hit rate cannot be treated as an 80% total-bill reduction.

  1. Keep the prefix stable

    Put durable instructions and repeated reference material first. Append the changing user request after that stable prefix instead of rebuilding the prompt in a different order.

  2. Read the usage fields

    Track prompt_cache_hit_tokens and prompt_cache_miss_tokens from API responses. A projected hit rate is not a billing fact until these fields show it.

  3. Price the observed mix

    Split measured input into hit and miss tokens, add output tokens, and calculate cost per accepted result using each pricing period's rates. Reprice when prompt structure or traffic cadence changes.

DeepSeek is cheaper than the API shortlist, before quality adjustment

DeepSeek wins the raw-token comparison against these three models. It does not automatically win the business outcome because these models differ in capability, tools, latency, and failure rate.

Use one normalized workload to make the sticker prices comparable: 1M total input tokens plus 250K output tokens, with no cache discount. That workload costs $0.30 off-peak or $0.60 peak on Flash, and $1.155 or $2.31 on Pro.

OpenAI GPT-5.6 Sol currently charges $4 input and $20 output per million tokens at standard short-context rates, with promotional pricing available at least through November 21, 2026. The normalized workload costs $9.00 when each request stays at or below OpenAI's long-context threshold, or 7.79x off-peak Pro and 3.90x peak Pro. Prompts above 272K input tokens trigger 2x input and 1.5x output pricing for the full request, and cache writes cost 1.25x uncached input.

Claude Sonnet 5 charges $2 input and $10 output per million tokens. The same workload costs $4.50, or 3.90x off-peak Pro and 1.95x peak Pro. Anthropic includes the 1M context window at standard pricing and offers a 50% Batch discount, so latency-tolerant work can narrow the gap.

Gemini 3.1 Pro Preview charges $2 input and $12 output per million tokens when a prompt stays at or below 200K. The normalized aggregate workload costs $5.00 if every request stays under that boundary, or 4.33x off-peak Pro and 2.16x peak Pro. Put the same workload into one request above 200K and Google's rates rise to $4 input and $18 output, producing an $8.50 bill, or 7.36x off-peak Pro and 3.68x peak Pro.

Historical August chart of normalized API costs across DeepSeek, Claude, Gemini, and OpenAI
August 2026 illustration of 1M input plus 250K output; the current DeepSeek and OpenAI totals are stated above

The price spread is large, but the dollar spread is smaller than the multiples make it sound. Moving the normalized workload from Pro to GPT-5.6 Sol adds $7.845 against off-peak Pro or $6.69 against peak Pro. One human correction can erase that saving, while a reliable high-volume route can compound it.

Use the DeepSeek vs ChatGPT comparison when model quality, product features, and data handling matter more than the token table. Use the cheapest AI API shortlist when the job can accept smaller budget models and gateways. A pricing page can normalize rates; only your evaluation can normalize outcomes.

The hidden budget risk is repricing, not an annual lock

DeepSeek's peak/off-peak schedule makes execution time a budget input. Its live pricing page also reserves the right to adjust prices and recommends topping up for actual usage. It no longer carries the old warning of an undated, significant upcoming increase.

That changes the procurement decision. Store the pricing-page verification date beside every budget model, make the rate and execution period editable inputs, and define the price change that would trigger a reroute.

There is no monthly or annual API trap

DeepSeek lists token-based deduction, not a monthly seat or annual self-serve contract. Topped-up balance does not expire, according to its current Help Center. The API uses granted balance first when both granted and topped-up balances exist.

Unused balance is refundable through Billing > Refunds for online payments, according to DeepSeek's current Help Center; corporate bank-transfer refunds require a ticket. Its Open Platform terms qualify that promise: requests are reviewed, approved refunds cover the remaining unspent amount less necessary costs, and partial refunds or refunds of spent usage are not supported. That makes a modest top-up easier to manage than a large prepaid balance.

When balance is insufficient, the API returns HTTP 402 and tells you to top up. The operational risk is interruption. Set a balance alert before production traffic finds that boundary for you.

Integration and concurrency can cost more than tokens

Flash lists 2,500 concurrent requests and Pro lists 500. Both now support the Responses API, while only Flash supports vision. A buyer who ignores those differences may pay in queueing, fallback logic, or integration work even when the token bill looks modest.

The cheapest safe architecture keeps routing reversible. Use one internal interface, record the chosen model, token usage and pricing period, and preserve a tested fallback. Editable rates turn a pricing change into a budget adjustment rather than an application rewrite.

DeepSeek's price history split Flash from Pro

DeepSeek's current tiers do not make every line item uniformly cheaper than its historic V3 schedule. Flash remains the volume-pricing route, but peak hours and Pro's higher rates change the comparison.

DeepSeek's V3 pricing announcement listed $0.07 cache-hit input, $0.27 cache-miss input, and $1.10 output per million tokens. Against those historic rates, today's off-peak Flash is 95.71% cheaper on cache-hit input, 44.44% cheaper on cache-miss input, and 45.45% cheaper on output. At peak hours, Flash cache hits are still 91.43% cheaper, but cache-miss input is 11.11% higher and output is 9.09% higher.

Pro moves differently. Off-peak cache-hit input is 68.57% cheaper than historic V3, but its $0.66 cache-miss input is 144.44% higher and $1.98 output is 80.00% higher. Peak Pro cache hits are 37.14% cheaper, while cache-miss input is 388.89% higher and output is 260.00% higher.

That history explains the current decision rule. Flash is the volume-pricing play. Pro asks the workload to earn a higher rate through better outcomes. Scheduling and future repricing can change both, so the verification date and time band belong next to the arithmetic.

The Monday move: measure one route before moving the budget

On Monday, take 100 sanitized requests from one bounded workflow and run a one-week routing test. Do not start with an enterprise-wide migration.

  1. Define acceptance first

    Choose a result a reviewer can score consistently: valid schema, correct classification, passing tests, or an approved support draft. Remove customer, personal, and confidential data from the evaluation set.

  2. Send the base load to Flash

    Run all 100 requests through deepseek-flash. Record input, output, cache-hit, cache-miss, retry, latency, acceptance, and peak/off-peak data for each request.

  3. Escalate only failures

    Send failed or ambiguous cases to deepseek-v4-pro. Measure whether Pro clears the quality floor and reduces failed work enough to earn its workload-specific premium. For the uncached support mix, that premium is 3.85x before human review.

  4. Set two budget triggers

    Keep Flash as the default only while it clears the quality floor. Keep DeepSeek as the provider only while the live normalized cost stays below your tested fallback after any price change.

The output of this test is one decision-ready number: cost per accepted result. It converts DeepSeek's token rates into a workflow budget finance can revisit without rebuilding the analysis.

Frequently asked questions

Is DeepSeek paid or free?

DeepSeek's official consumer app is free, with no ads or in-app purchases. API access is separate and charges peak or off-peak rates for cache-hit input, cache-miss input, and output tokens.

How much does DeepSeek V4 Pro cost?

deepseek-v4-pro costs $0.022 per million cache-hit input tokens, $0.66 per million cache-miss input tokens, and $1.98 per million output tokens off-peak. Peak rates are $0.044, $1.32, and $3.96 respectively. It serves DeepSeek-V4-Pro-0813, with a 1M context window, 384K maximum output, and a listed concurrency limit of 500.

How much is DeepSeek per month?

Consumer chat costs $0 per month. The API has no fixed monthly price: split monthly cache-hit input, cache-miss input, and output tokens by peak/off-peak period, then multiply each total by its per-million rate.

How is DeepSeek cheaper than ChatGPT?

On a price-only API workload of 1M input plus 250K output tokens, DeepSeek Pro costs $1.155 off-peak or $2.31 peak, while OpenAI GPT-5.6 Sol costs $9.00 at its current standard short-context rate. That is a 7.79x or 3.90x difference. Compare accepted outcomes because model quality and tool support differ; these are API costs, not ChatGPT subscription prices.

Does DeepSeek offer a student discount?

DeepSeek does not list a student-specific plan or discount on the official pricing or app pages verified on October 4, 2026. Students using the consumer web or mobile app already pay $0.

Can I get a refund from DeepSeek?

Unused API balance can be refunded. For online payments, use Billing > Refunds; corporate bank-transfer payments require a support ticket. Topped-up balance does not expire. Refunds are subject to review and necessary costs under the Open Platform terms; partial refunds and refunds of spent usage are not supported.

Did DeepSeek pricing change in 2026?

Yes. Peak/off-peak pricing took effect on August 16, 2026, at 16:00 UTC, and V4.1 Flash brought another pricing adjustment in September. Current peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday, excluding Chinese public holidays. All other hours are off-peak, at half the peak rates.

Should I choose V4 Flash or V4 Pro?

Choose the current Flash model, deepseek-flash, for the bounded base load and Pro for cases where it demonstrably improves accepted results. Pro's uncached premium is 4.4x on input and 3.3x on output in the same pricing period. At the support example's fixed token mix it is 3.85x, requiring 74.03% less comparable billed work to break even on model spend.

Get the AI Tools Map for Business Owners

The AI Tools Map for Business Owners turns model prices into a practical adoption stack, with the quality floor, routing role, and cost trigger for each tool. Subscribe to get the next edition free.

Last Updated
Category
AI
Related Articles
Gamma Alternatives in 2026: Beautiful.ai, Prezi and More

Gamma Alternatives in 2026: Beautiful.ai, Prezi and More

Compare Gamma alternatives for brand control, editable PowerPoint and Google Slides, pitch decks and live talks. Current plans and clear tradeoffs.Oct 11, 2026AI
Reflection AI Beam: What It Means for Coding and Agents

Reflection AI Beam: What It Means for Coding and Agents

Reflection AI Beam explained: October release plans, vendor benchmark claims, beta access and whether builders should wait before switching models.Oct 10, 2026AI
Claude for Google Workspace

Claude for Google Workspace

How to install and deploy Claude in Docs, Sheets and Slides, control edits, and test the paid-plan beta with three practical team workflows.Oct 10, 2026AI
Claude Dashboards

Claude Dashboards

Set up Claude Dashboards beta, inspect SQL and refresh timestamps, check a sales metric, and decide when to keep your BI tool.Oct 10, 2026AI
Granola AI Review

Granola AI Review

Granola reviewed for business meetings: bot-free notes, current pricing, templates, recipes, consent, platform limits and who should choose an alternative.Oct 10, 2026AI
Fyxer Review

Fyxer Review

Fyxer reviewed through current pricing, help docs and public feedback: inbox sorting, drafts, meetings, privacy and who should pay for Pro or Team.Oct 9, 2026AI
GPT 6 Explained: Models, Release Dates, Access, and Prices

GPT 6 Explained: Models, Release Dates, Access, and Prices

GPT-6 Astra, Sol, 6.1 Sol and Luna explained: release dates, ChatGPT plan access, Work and Codex availability, and verified API costs.Oct 9, 2026AI
Is ChatGPT Free? What You Get for $0

Is ChatGPT Free? What You Get for $0

Yes. See what ChatGPT Free includes, its work limits, ads, account access, and when the $8 Go or $20 Plus plan is worth paying for.Oct 9, 2026AI
Newsletter

One letter, every Sunday.Working systems, not hot takes.

Weekly. No spam. Unsubscribe anytime.