Together AI Pricing (2026): PTUs Buy Capacity, Not Savings
Together AI pricing, verified August 2026: serverless token rates, $0.05 PTUs, GPU costs, hidden fees, break-evens, and alternatives.
- TTogether AI
- FFireworks AI
- GGroq
- OOpenRouter

Together AI's cheapest path is serverless, and its newest $0.05-per-PTU-minute option is not a list-price discount against its own token rates. One continuously reserved PTU costs $2,160 in a 30-day month; buy it for guaranteed capacity and a 99% uptime SLA, not for cheaper idle tokens.
Together AI pricing at a glance
Together AI is best bought in stages: start with prepaid serverless usage, move steady and reliability-sensitive traffic to Provisioned Throughput, and reserve dedicated hardware only when the workload needs single-tenant control or custom models. The expensive mistake is treating those stages as volume discounts. They meter different things and transfer different operating risks.

Together AI does not sell a simple Free, Pro, and Business ladder. It sells serverless tokens and media outputs, reserved token capacity, managed dedicated GPUs, raw GPU clusters, sandbox compute, storage, and fine-tuning. A single product can touch several meters. An agent may pay for model tokens, sandbox CPU and memory, and a fine-tuned endpoint that keeps billing between requests.
Every price and limit on this page was verified against Together AI's live pricing page, billing documentation, and product documentation on August 22, 2026. That date matters. The current page already differs materially from June pricing coverage: it lists MiniMax M3, Kimi K3, GLM-5.2, lower live dedicated hardware rates, and the new Provisioned Throughput layer.
The first decision rule is simple. Use serverless when demand is uncertain or bursty. Consider a PTU when a bounded model carries a steady base load and reserved capacity has business value. Choose Dedicated Inference when you need a custom model, single-tenant hardware, or configuration control. Choose a GPU cluster only when your company wants to operate the underlying compute rather than buy a managed inference outcome.

That distinction makes Together AI attractive for a funded founder who wants one API path from experiment to reserved infrastructure. It is less attractive for a solo builder chasing the lowest spot token rate, or for a team that cannot forecast even one month of demand. Flexibility is the serverless product. Commitment is a separate purchase.
Serverless inference is the default, and model choice dominates the bill
Serverless is the right starting point because there is no provisioning cost and no usage minimum beyond the $5 platform-access purchase. You send requests to a shared fleet and pay for the work completed. The model, token direction, cache rate, and output length matter far more than the Together AI brand name on the invoice.
The current chat catalog spans a wide range. LFM2.5-8B-A1B costs $0.03 per million input tokens and $0.12 per million output tokens. DeepSeek V4 Flash 0731 costs $0.14 input, $0.03 cached input, and $0.28 output. gpt-oss-120B costs $0.15 input and $0.60 output. At the upper end of the displayed text list, Kimi K3 costs $3 input, $0.30 cached input, and $15 output per million tokens.
That Together-hosted DeepSeek rate is a provider-specific meter. The separate DeepSeek pricing breakdown covers the direct model economics; use Together's live page for a Together invoice.
That is a 125x spread between the $0.12 LFM output rate and the $15 Kimi K3 output rate. Optimizing a provider while ignoring model routing is like negotiating a small discount on freight while sending every parcel by air. The first cost control is to assign the cheapest model that clears the quality floor for each task.
A useful agent workload makes this concrete. Suppose 1,000 calls each consume 8,000 input tokens and produce 1,000 output tokens. That is 8 million input tokens and 1 million output tokens:
- DeepSeek V4 Flash 0731 costs $1.40 total, or $0.0014 per call.
- gpt-oss-120B costs $1.80 total, or $0.0018 per call.
- MiniMax M3 costs $3.60 total, or $0.0036 per call.
- GLM-5.2 costs $15.60 total, or $0.0156 per call.
- Kimi K3 costs $39 total, or $0.039 per call.
The largest difference is not provider overhead. It is the chosen model and the amount of expensive output it produces. A customer-support classifier that returns one label does not need the same output budget as a research agent drafting a long response. Put a hard output cap on the bounded job before shopping for a cheaper endpoint.
This is why a broader cheapest AI API comparison is useful only after the job is normalized. Comparing providers on different prompt lengths, response lengths, cache assumptions, or models produces a neat number that cannot guide a budget.
Cached input can cut the expensive repeated prefix
Cached input is the strongest self-serve discount when prompts repeat. It means Together recognizes an unchanged prefix, such as a long system instruction or stable reference context, and bills those tokens at a lower rate on supported models.
For the same 1,000-call workload, assume 80% of the 8 million input tokens qualify for cached pricing. MiniMax M3 falls from $3.60 to $2.064. GLM-5.2 falls from $15.60 to $8.304. Kimi K3 falls from $39 to $21.72. Those reductions are 42.7%, 46.8%, and 44.3%, respectively.
The operational consequence is bigger than the discount percentage. A senior operator should separate the stable prefix from request-specific context, record cached and uncached tokens independently, and watch for prompt edits that invalidate the cache. A prompt rewrite can raise spend without changing traffic, which looks like a pricing problem until the cache data explains it.
Batch is cheaper only when the model and workflow qualify
Together AI's Batch API gives selected serverless models up to a 50% discount in exchange for asynchronous completion. It is the right meter for offline classification, synthetic data, evaluations, and bulk summaries. It is wrong for a live chat, checkout assistant, or customer-service flow waiting on the next response.
The current batch contract allows up to 50,000 requests in a 100 MB input file and up to 30 billion enqueued tokens per model. The completion window is fixed at 24 hours and is a best-effort target. Together recommends batches of 1,000 to 10,000 requests. Successful responses are billed, failed requests are not, and canceling a batch does not erase the cost of responses already completed.
The 50% headline does not apply universally. Only selected serverless models receive the discounted rate, and some models cannot run through Batch at all. Dedicated endpoints can accept batch jobs, but the discount does not apply there. Check the live batch model list before building a forecast.
For a current discounted example, 8 million input plus 1 million output tokens on Llama 3.3 70B cost $9.36 at its $1.04 input and output rates. A qualifying 50%-off batch costs $4.68. The saving is meaningful because the job accepts delayed completion, not because every serverless request became cheaper.
Images, video, audio, and storage use different units
Serverless does not stop at text. The live image catalog ranges from $0.0017 to $0.134 per image or megapixel across the displayed models. The video list ranges from $0.115 to $3.20 per generated video. Speech-to-text ranges from $0.0015 to $0.0045 per audio minute, while text-to-speech ranges from $4 to $65 per million characters. Multilingual e5 large instruct embeddings cost $0.02 per million tokens, and Llama Guard 4 12B moderation costs $0.20 per million tokens.
Those ranges are not interchangeable. Together notes that displayed image and video prices use the lowest listed resolution or duration, and that exceeding default generation steps can add cost. A product manager budgeting 100,000 short video generations needs the exact model, duration, resolution, audio setting, and retry rate. Multiplying the cheapest thumbnail price by the production count will understate the line item.
Serverless is still the right default for most new workloads. Its wall is not the token rate. Its wall is shared capacity, dynamic rate limits, and the absence of a public uptime commitment equivalent to the Provisioned product. Once that becomes a customer-facing risk, the next decision is capacity, not a generic upgrade.
Provisioned Throughput turns capacity into a monthly commitment
Provisioned Throughput is worth buying only when reserved token capacity solves a business problem. Together AI launched it on July 8, 2026 as a middle layer between best-effort serverless and fully configurable Dedicated Inference. It carries a 99% uptime SLA, a one-month minimum, and a public list price of $0.05 per Provisioned Throughput Unit, or PTU, per minute.
A PTU is not a token bundle. It is a reserved rate of tokens per minute for one supported model, held for the customer. Input, cached input, and output consume that capacity at different rates. The current live page supports three public schedules:
MiniMax M3 delivers 166,667 input tokens, 833,333 cached input tokens, or 41,667 output tokens per minute per PTU. Kimi K3 delivers 16,667 input, 166,667 cached input, or 3,333 output tokens. GLM-5.2 delivers 35,714 input, 192,308 cached input, or 11,364 output tokens. A mixed request consumes a blend of those capacities.
At $0.05 per minute, one PTU costs $3 per hour, $72 per day, and $2,160 over a 30-day month. Together's calculator footnote uses about 43,800 minutes for continuous average-month provisioning, which produces about $2,190. Finance should use the exact contract and calendar month rather than assuming every month is a flat $2,160.
The PTU occupancy test
The revealing calculation is what one PTU would have cost on Together's own serverless meter. Devote a MiniMax M3 PTU to input for 30 days and it can deliver about 7.2 billion input tokens. At the serverless rate of $0.30 per million, that capacity is worth about $2,160. Devote it to output and it can deliver about 1.8 billion output tokens. At $1.20 per million, that is also about $2,160. The cached-input route lands at the same point within rounding.
Kimi K3 and GLM-5.2 are calibrated the same way. Public PTU capacity reaches parity with public serverless pricing at essentially full occupancy over a 30-day month. The published calculator can show large savings against a selected commercial model, but it explicitly compares PTU spend with that outside model's list price. It does not establish a blanket discount against Together's own serverless rates.
This is the PTU occupancy test: capacity value equals serverless value only when the reserved unit is fully used. At 50% utilization, the effective MiniMax M3 rates double to $0.60 per million input tokens, $0.12 per million cached input tokens, or $2.40 per million output tokens. At 25%, they quadruple to $1.20, $0.24, and $4.80. At 70%, the effective rate is about 1.43x serverless, roughly $0.43 input and $1.71 output.

That does not make PTUs a bad product. It clarifies the product. A mid-market CTO running a customer-facing agent may rationally accept 20% idle capacity if the reserved lane reduces throttling risk and the 99% uptime SLA has contractual value. A founder with weekday-only demand should not pay for nights and weekends unless the capacity guarantee protects enough revenue to cover the premium.
The public page says higher commitment levels can receive discounts but does not publish the schedule. A negotiated discount moves the break-even below 100% occupancy. Until sales puts that rate in an order form, model the public list price and treat any discount as upside.
A 99% SLA still needs business interpretation
A 99% monthly uptime target permits roughly 7 hours and 18 minutes of downtime in a 30.4-day month. That is stronger than best-effort access, but Groq's enterprise Performance tier publishes 99.9%, so 99% is not the strongest commitment in this shortlist. The business should compare the written remedy, maintenance exclusions, regional design, and fallback route, not stop at the presence of three characters and a percent sign.
Provisioned Throughput is most defensible when the application has a stable base load, predictable model choice, and a customer promise that shared-fleet throttling can break. It is less defensible when traffic swings by 10x, the model changes weekly, or offline work can move to Batch. In those cases the reservation converts uncertainty into paid idle time.
The choice changes the Monday budget conversation
Before July 8, a team choosing Together had a broad jump from shared per-token serverless to dedicated GPU infrastructure. PTUs add a middle budget line: reserved model throughput without taking on GPU sizing or a custom serving stack. That lets finance approve a fixed monthly capacity envelope while engineering keeps the same inference API.
The consequence is a new operating metric. Token spend alone is no longer enough. The owner needs peak requests per second, input and output token mix, cache hit rate, capacity occupancy, throttled requests, and accepted outcomes. If the dashboard cannot show those six measures, the company cannot tell whether a PTU is protection or waste.
One MiniMax M3 PTU buys about 600,000 worked agent calls
One PTU can support about 600,000 MiniMax M3 calls in a 30-day month when each call uses 8,000 uncached input tokens and produces 1,000 output tokens. That is the same worked workload priced earlier at $0.0036 per serverless call, so 600,000 serverless calls cost $2,160. The capacity and token meters meet at full occupancy by design.
The capacity calculation uses both sides of the request. Eight thousand input tokens consume about 0.048 PTU-minutes at MiniMax M3's 166,667 input-token-per-minute schedule. One thousand output tokens consume another 0.024 PTU-minutes at 41,667 output tokens per minute. The complete request therefore consumes about 0.072 PTU-minutes. Divide the month's 43,200 PTU-minutes by that request footprint and the result is about 600,000 calls.
This creates a decision-ready operating budget:
At 80% occupancy, the PTU supports about 480,000 calls. The equivalent serverless bill is $1,728, leaving a $432 monthly reservation premium. The PTU is 25% more expensive than the consumed serverless work. A customer-facing product can accept that premium if reserved capacity and the 99% SLA protect more than $432 of revenue, support time, or reputational cost.
At 70% occupancy, it supports about 420,000 calls. Serverless costs $1,512, so the capacity premium is $648. At 50%, it supports 300,000 calls and serverless costs $1,080, leaving half the PTU bill as paid idle capacity.
Cache changes how many calls fit inside the unit. If 80% of the input is cached, each request uses about 0.04128 PTU-minutes instead of 0.072. One PTU can then carry roughly 1.0465 million of the same calls at full occupancy. Serverless pricing falls at the same time, from $0.0036 to $0.002064 per call, so the full-load cost still converges near $2,160.
The business gain from cache is therefore throughput as well as unit cost. A long, stable system prefix lets the same reserved capacity carry more customer work. A prompt change that destroys cache reuse reduces PTU headroom and raises serverless spend simultaneously.
Now scale the decision to a budget line. Ten PTUs cost $21,600 for a 30-day month and carry about 6 million of the uncached worked calls at full occupancy. At 70%, they carry 4.2 million calls. The equivalent serverless bill is $15,120, leaving a $6,480 monthly premium for reserved capacity and the SLA.
That $6,480 is the number a CTO should defend. If shared-fleet throttling threatens more than $6,480 of gross margin, support burden, or contractual penalties, the reservation can earn its place. If it does not, serverless keeps the same $6,480 available for product work. A generic "we are spending $15,000 on tokens" argument never reaches this decision.
Dedicated Inference does not have an equally clean public token break-even because output throughput depends on the model, quantization, batch shape, hardware count, and configuration. Any article that divides a GPU's monthly price by one generic token rate and declares a universal crossover is hiding the variable that decides the answer. Measure the candidate model on the intended endpoint, then compare cost per accepted result at the observed utilization.
Dedicated Inference and GPU clusters solve different problems
Dedicated Inference is the right move when control matters more than self-serve flexibility. Together allocates single-tenant hardware, supports custom models, and offers autoscaling and traffic-spike handling. The live price is $5.49 per GPU-hour for an NVIDIA HGX H100 and $8.99 for an HGX B200. H200, B300, GB200 NVL72, and GB300 NVL72 prices require sales contact.
At continuous 30-day use, one H100 costs $3,952.80 and one B200 costs $6,472.80. Those are per-GPU figures before any multi-GPU requirement. The endpoint bills while it is running regardless of request volume, so an idle deployment is still a live budget item.
Do not use older dedicated-endpoint documentation to forecast the hardware bill. The docs currently expose H100, H200, and B200 numbers that conflict with the live pricing page. The current pricing page is the newer source and is the one verified here. This is exactly why an infrastructure price should carry a date.
Dedicated Inference earns the bill when a fine-tuned or uploaded model cannot run on shared serverless, when p99 latency needs stable hardware, when model serving requires custom settings, or when steady high utilization makes a reserved GPU economically superior. It loses when the endpoint spends much of the day waiting.
GPU clusters are rawer and can look cheaper for the wrong reason
GPU clusters hand more of the stack to the buyer. The current on-demand rates are $3.99 per GPU-hour for H100, $5.99 for H200, and $8.19 for B200. A 30-day H100 equivalent is $2,872.80, but that is not a cheaper substitute for the $3,952.80 Dedicated H100. The cluster buyer owns more deployment, serving, orchestration, observability, and utilization work.
Reserved cluster rates decrease with commitment. H100 falls to $3.69 for 7-30 days, $3.45 for 31-90 days, and $3.19 for 91-180 days. H200 falls from $4.99 to $4.15 to $3.99 across the same bands. B200 falls from $7.99 to $7.79 to $6.79. Terms of 181 days or longer require a quote.
The discount comes with hard edges. Together's GPU billing documentation says the full reservation is charged upfront or deducted once provisioned. Reservations are non-refundable, cannot be partially refunded, and cannot be paused. If the cluster scales above reserved capacity, the extra capacity is billed at on-demand rates.
For a 91-180-day H100 reservation, the normalized 720-hour month is $2,296.80 per GPU. That is $576 below the on-demand equivalent. The saving is valuable only if the GPU remains usefully occupied. A four-GPU reservation left half idle does not become economical because the hourly line is lower.
Storage also has its own lifecycle. Shared storage costs $0.16 per GiB-month, making 1 TiB about $163.84 per month. The volume persists and continues billing after a cluster is terminated until someone deletes it. That is operationally useful for recreating compute around durable data, but it also turns abandoned volumes into silent spend.
Fine-tuning can be cheap to train and expensive to keep alive
Fine-tuning costs two separate things: the training job and the model's later deployment. The training line can look tiny. Hosting the result can become the dominant monthly expense.
Together calculates standard training from the dataset tokens multiplied by epochs, plus optional evaluation tokens multiplied by evaluation runs. For models up to 16B parameters, supervised fine-tuning costs $0.48 per million processed tokens with LoRA and $0.54 with full fine-tuning. Direct Preference Optimization, or DPO, costs $1.20 with LoRA and $1.35 with full tuning.
For 17B-69B models, the four rates are $1.50, $1.65, $3.75, and $4.12. For 70B-100B models, they are $2.90, $3.20, $7.25, and $8. The standard minimum job charge is $4.
Larger and named models use specialized pricing. DeepSeek V4 Flash LoRA is $6 per million processed tokens for supervised fine-tuning and $15 for DPO, with a $12 minimum. gpt-oss-120B is $5 and $12.50 with a $6 minimum. Kimi K2.6 or K2.7-Code is $15 and $37.50 with a $60 minimum. GLM-5.2 is $40 and $100 with a $60 minimum.
That model-specific spread changes a seemingly simple experiment. A 20-million-token dataset run for three epochs processes 60 million tokens. On a standard up-to-16B LoRA job, supervised training costs $28.80. On specialized GLM-5.2 LoRA, the same processed-token count costs $2,400. Dataset size did not change; the base-model price did.
The training job is still not the lifecycle cost. A fine-tuned model requires deployment capacity and sufficient prepaid balance. A single current Dedicated H100 left running for 30 days costs $3,952.80. A team can spend $28.80 creating a small adapter and then spend more than 137x that amount each month keeping one H100 active.
This is where a prototype budget often breaks. The data scientist reports the training run. Finance later sees the serving bill. Approve the experiment with a deployment plan attached: expected hours online, GPU count, autoscaling floor, request volume, and shutdown owner.
Canceled training has a narrower refund than many buyers expect. Together charges the completed steps and refunds only the uncompleted work. Standard platform fees are otherwise non-refundable by default under the terms. Treat a training job as consumed compute, not a reversible software subscription.
Sandbox and Code Interpreter are separate agent meters
Code Sandbox costs $0.0446 per vCPU-hour plus $0.0149 per GiB of RAM per hour. A sandbox running 2 vCPUs and 8 GiB for 160 hours costs about $33.34. Code Interpreter uses a different meter at $0.03 per 60-minute session.
For an agent product, those charges sit beside model inference. If 10,000 monthly agent runs each open a one-hour interpreter session, the interpreter line alone is $300 before model tokens. If sessions can be reused safely, the product architecture changes cost per completed job. If they remain open after work ends, idle execution becomes the equivalent of idle GPU capacity at a smaller scale.
The right unit is therefore cost per accepted outcome, not cost per million tokens. The worked MiniMax M3 call costs $0.0036; add one $0.03 Code Interpreter session and the path costs at least $0.0336 before retries or other tool calls. Instrument the full path.
Together AI is not free, even when a model shows $0 tokens
Together AI has no current free trial and requires a minimum $5 prepaid credit purchase before platform access. That makes the free-plan answer unambiguous: there is no zero-payment free account for API use.
The live catalog does list Ternary Bonsai 27B at $0 per million input and $0 per million output tokens. That can make ongoing inference cost $0 while the listing remains available, but the account still needs the $5 access purchase. The $5 is not described as a monthly subscription. Purchased prepaid credits currently do not expire.
Who never needs to pay beyond the first $5? A hobbyist or evaluator who uses only the current zero-priced model, stays within rate limits, and never starts a paid service may leave the purchased balance untouched. That is a narrow case, not a production free tier. Model availability, capacity, and pricing can change, and a product should not promise customers that a promotional or zero-priced endpoint will remain its permanent backend.
Free signup credits offered in the past have ended. Together's current support policy says users must buy at least $5 and connect a payment method. There is no negative balance for ordinary prepaid access; when the purchased balance reaches zero, API access is suspended until more credit is added. Contract customers follow their order terms.
Credits do not expire, but they are not a refundable savings account
Purchased credits currently have no expiration date. That is friendlier than an annual-expiry policy, but it does not make overfunding harmless. Together's terms say fees paid are non-refundable by default, payment obligations are non-cancelable, and partial months are not prorated unless an order form says otherwise.
Auto-recharge deserves careful configuration. The billing documentation lists a $25 default recharge amount and warns that setting a threshold above the current balance can trigger immediate repeated purchases until the gap is filled. Use a threshold tied to expected burn, alert on each purchase, and keep a project-level budget outside the credit balance.
Build Tiers raise limits through lifetime spend
Together's five Build Tiers are rate-limit bands based on lifetime spend, not monthly feature plans. Tier 1 starts at $5 and lists 600 LLM requests per minute. Tier 2 starts at $50 and lists 1,800. Tier 3 starts at $100 and lists 3,000. Tier 4 starts at $250 and lists 4,500. Tier 5 starts at $1,000 and lists 6,000.
Embedding limits rise from 3,000 RPM at Tier 1 to 10,000 at Tiers 4 and 5. The rerank allowance rises from 500,000 to 5,000,000 across the five tiers. These are not universal promises for every model. Together can apply model-specific restrictions, and its current serverless documentation also describes dynamic limits that react to live capacity and recent successful traffic.
That combination matters for a launch. A Tier 5 account can still see a 429 response if a popular model has tighter capacity or if traffic jumps far above its recent pattern. Serverless solves provisioning, not every burst. Batch, PTUs, or Dedicated Inference are the escape paths, each with a different cost and latency tradeoff.
The hidden costs are mostly idle meters and policy edges
Together AI's visible token prices are generally clear. The expensive surprises sit between products, in commitments, idle resources, and assumptions that look universal but are not.
PTU calendar arithmetic: $0.05 per minute produces $2,160 in a 30-day month. The live calculator footnote uses about 43,800 minutes, which produces about $2,190. That $30 difference per PTU becomes $3,000 across 100 PTUs. Use the contract's billing convention.
PTU underuse: a public PTU is calibrated to Together serverless value at full 30-day occupancy. At 50% occupancy, effective token cost doubles. The product can still earn that premium through capacity and SLA value, but the premium should be visible.
Dedicated idle time: Dedicated Inference bills the allocated hardware while it runs, even with zero requests. Fine-tuned model hosting belongs in the monthly forecast, not only the experiment ticket.
Non-refundable reservations: GPU cluster reservations are charged for the full term, cannot be paused, and cannot be partially refunded. A forecast error becomes committed spend.
On-demand burst overage: scaling beyond reserved GPU capacity is billed at the higher on-demand rate. A stable baseline can still produce a surprise if burst assumptions are too low.
Persistent storage: storage survives cluster termination and keeps billing until deleted. Every cluster teardown needs an explicit keep-or-delete decision.
Batch eligibility: the discount is model-specific, the 24-hour window is best effort, and completed responses remain billable after cancellation. Do not apply 50% to an entire offline budget without checking the live list.
Multimedia defaults: image and video prices can refer to the lowest resolution, duration, or default step count. Production quality settings and retries move cost per usable asset.
Auto-recharge behavior: an aggressive threshold can cause immediate purchases, and ordinary fees are non-refundable. Treat auto-recharge as a production continuity control with alerts, not a set-and-forget convenience.
Output drift: a prompt or model update that produces longer responses raises spend even at stable request volume. Track input, cached input, output, and accepted results separately.
Rate-limit drift: serverless limits can be dynamic and model-specific. Buying lifetime-spend Tier 5 does not reserve the underlying fleet.
There is no standard annual self-serve discount to model. Serverless is prepaid usage, PTUs carry a one-month minimum, GPU reservations publish 7-30, 31-90, and 91-180-day bands, and longer commitments go to sales. The annual-lock risk lives in a custom order form or a chain of infrastructure reservations, not in a public annual app plan.
Together AI alternatives: normalize the same model before choosing
Together AI is not uniquely cheap on every shared model. For gpt-oss-120B, its public serverless price exactly matches Fireworks AI and Groq at $0.15 per million input tokens and $0.60 per million output tokens. A job with 1 million input and 250,000 output tokens costs $0.30 on each provider before any cached-input advantage or account-specific agreement.
That parity is useful. It removes token price as the deciding factor and forces the buyer to compare cache pricing, rate limits, latency, provider control, fallback behavior, deployment paths, and SLA packaging.
Fireworks AI matches the base rate and adds a cheaper cache line
Fireworks AI lists gpt-oss-120B Standard serverless at $0.15 input, $0.015 cached input, and $0.60 output per million tokens. The uncached normalized job is the same $0.30 as Together. Fireworks also advertises $1 in starting credits.

Fireworks wins this narrow comparison when the workload has a large reusable prefix because its current gpt-oss cached-input rate is explicit and low. It also offers Standard, Priority, and Fast serverless tiers, which create another latency-price choice. The limitation is that those extra tiers complicate a supposedly simple rate comparison, and its dedicated GPU price is a different product from Together's PTU.
Choose Fireworks when cache-heavy open-model serverless is the main need and its service-tier controls fit the application. Choose Together when the path from the same API into PTUs, dedicated model inference, fine-tuning, sandboxes, or raw clusters is the larger strategic value.
Groq sells speed at the same token rate
Groq lists gpt-oss-120B at the same $0.15 input and $0.60 output rates, so the normalized job is also $0.30. Its current model page lists about 500 tokens per second, a Developer-plan limit of 250,000 tokens per minute, and 1,000 requests per minute for this model.

Groq is the sharper shortlist entry when low generation latency is the product promise. Its enterprise Performance tier adds provisioned throughput with a 99.9% availability SLA, but public pricing is not listed. That makes self-serve token cost easy to compare and reserved-capacity cost impossible to normalize without a quote.
Choose Groq for an interactive path where its supported model set and speed matter more than Together's broader training and infrastructure stack. Skip it when the workload needs Together's fine-tuning, GPU cluster, sandbox, or public PTU path under one account.
OpenRouter can be cheaper, with an aggregator tradeoff
OpenRouter currently lists gpt-oss-120B at $0.03 input, $0.03 cached input, and $0.17 output per million tokens. The normalized 1-million-input and 250,000-output job costs $0.0725 before credit-purchase fees. Allocating its 5.5% fee across fully used credits brings that to about $0.0765, although the $0.80 minimum fee dominates small top-ups.

OpenRouter wins the spot-price example by routing across upstream providers. It also provides automatic fallback and a much broader model catalog. That is not the same product as reserving Together capacity or choosing a single infrastructure provider. Route selection, data policy, latency, upstream availability, and provider changes become part of the architecture.
Its free tier is more accessible than Together's: the current pricing page lists more than 25 free models and 50 requests per day. The tradeoffs include low free limits, a 5.5% pay-as-you-go platform fee, and the right to expire unused credits after one year. Together currently says purchased credits do not expire.
Choose OpenRouter when broad model access, fallbacks, and the current routed spot price matter most. Choose Together when direct provider control, a predictable path into reserved capacity, model training, or dedicated infrastructure matters more than the cheapest routed request.
The normalized result is blunt. For gpt-oss-120B, Together, Fireworks, and Groq tie on uncached token price. OpenRouter is cheaper in this snapshot, but it changes the supplier relationship. The choice flips on the operating requirement, not on a generic claim that one platform is cheaper.
Who should choose Together AI, and who should skip it
Together AI is the right platform for a company that expects its open-model workload to mature from uncertain calls into a deliberate capacity plan. It is not the automatic choice for every developer who wants a low-cost API.
A solo technical builder should start with serverless and keep the first purchase at $5. Route a bounded task to DeepSeek V4 Flash, gpt-oss-120B, or another model that clears the quality floor. Stay off PTUs and dedicated hardware until production logs prove steady demand. If the only goal is broad model sampling or free access, OpenRouter can be the cleaner first account.
A funded founder should use Together when the deployment ladder reduces future migration work. Serverless can validate the product, PTUs can reserve a stable open model behind the same API, and Dedicated Inference can host a fine-tuned or custom model. That continuity is worth more than a small token delta when the product roadmap already points toward reserved capacity.
A mid-market CTO should buy PTUs for a service commitment, not for a large invoice. Forecast occupancy, compare the 99% SLA with the customer promise, and value the capacity reservation against lost transactions or support escalation. If the application needs 99.9% availability, a 99% line is not enough by itself; design fallback or negotiate the contract.
A senior operator should move offline volume to Batch before reserving capacity. Classification, enrichment, evaluations, and synthetic-data jobs can receive up to 50% off on eligible models. That is a direct discount for accepting latency. A PTU is a capacity purchase and should come later.
A research or model team should use Dedicated Inference or GPU clusters only with an owner for utilization. Training, custom serving, and low-level control can justify the stack. A team that cannot monitor GPU-hours, idle replicas, burst usage, and persistent volumes is buying infrastructure it cannot govern.
Skip Together AI when one of four conditions holds. The cheapest routed token is the only priority; traffic is too erratic for a one-month reservation; the required model lives elsewhere; or the company needs a stronger public SLA without a sales negotiation. Fireworks, Groq, OpenRouter, or a direct model vendor may fit better.
- One account spans serverless, reserved token capacity, dedicated inference, fine-tuning, GPU clusters, sandbox compute, and storage.
- Current serverless rates are competitive on several open models, with substantial cached-input discounts on supported endpoints.
- Purchased prepaid credits currently do not expire.
- Provisioned Throughput adds a public $0.05/PTU-minute path between shared and dedicated inference.
- There is no zero-payment free trial; platform access requires a $5 purchase.
- Public PTU list pricing reaches parity with Together serverless only at full occupancy before negotiated discounts.
- A 99% PTU uptime SLA may be too weak for critical customer paths without fallback.
- Dedicated, reserved GPU, storage, and fine-tuned deployments can keep billing while request volume is zero.
- Live prices can conflict with older product documentation, so procurement needs a dated source.
The Monday move: measure one workload before reserving anything
On Monday, choose one bounded production workflow and build a seven-day cost record. The goal is not to test every model. It is to learn whether the load belongs on serverless, Batch, PTU, or dedicated capacity.
Define one accepted outcome
Pick a task with a scorable result: a valid classification, a passing code change, an approved support draft, or a completed extraction. Cost per accepted result is the business metric.
Record the full usage shape
Capture request count, uncached input, cached input, output, retries, peak requests per second, 429 responses, latency, sandbox time, and accepted outcomes for seven days. Separate weekday peaks from nights and weekends.
Price the serverless baseline
Apply the live model rates to the measured token mix. Move eligible offline work to Batch and recompute. Cap unnecessary output and preserve stable prompt prefixes so the baseline includes attainable cache savings.
Run the PTU occupancy test
Use the live PTU calculator with peak request rate, cache hit rate, and token mix. Compare the required PTU count with average use. At 80% occupancy, public list price carries a 25% effective token premium; decide whether reserved capacity and the 99% SLA earn it.
Escalate only for a named requirement
Move to Dedicated Inference for custom models, single-tenant control, or stable hardware performance. Move to GPU clusters only when the team owns the serving or training stack. Put the shutdown and storage-cleanup owner beside the budget owner.
The decision output is one sentence finance and engineering can share: "This workload stays serverless," "this offline lane moves to Batch," or "this customer path reserves N PTUs because the capacity guarantee is worth the measured idle premium." That is a Monday move, not an infrastructure migration based on a pricing headline.
Together AI pricing FAQ
How much do 1,000 tokens cost on Together AI?
Divide the model's per-million rate by 1,000 and price input and output separately. MiniMax M3 costs $0.0003 per 1,000 uncached input tokens and $0.0012 per 1,000 output tokens. Kimi K3 costs $0.003 input and $0.015 output per 1,000 tokens.
How much does Together AI cost per month?
Serverless has no fixed monthly subscription; after the $5 access purchase, monthly cost follows usage. One PTU costs $2,160 in a 30-day month or about $2,190 using the pricing page's 43,800-minute average-month assumption. One continuously running Dedicated H100 costs $3,952.80 over 720 hours at the current $5.49 hourly rate.
Is Together AI free?
No. Together AI does not currently offer a free trial and requires a minimum $5 prepaid credit purchase for platform access. The live catalog lists Ternary Bonsai 27B at $0 input and output token rates, but the account still requires the $5 purchase.
Does Together AI give free credits?
Not as a standing signup offer. Prior signup promotions have ended, and current access requires a $5 purchase. Together also operates an invite-only research credits program for student projects outside formal classes, but it is not widely available.
What are Together AI's free-tier limits?
There is no conventional free tier. A purchased $5 balance places the account in Build Tier 1, which currently lists 600 LLM requests per minute, 3,000 embedding requests per minute, and a 500,000 rerank limit. Model-specific and dynamic limits can be lower.
What are the best alternatives to Together AI?
Fireworks AI is the closest broad open-model inference and fine-tuning alternative, Groq is compelling for supported low-latency models, and OpenRouter is strongest for broad model routing and fallbacks. On gpt-oss-120B, Together, Fireworks, and Groq currently share the same $0.15 input and $0.60 output rate per million tokens; OpenRouter's routed price is lower in this snapshot but adds a credit-purchase fee and an aggregator layer.
Does Together AI offer a student discount?
Together AI does not publish a standing student pricing tier. Its invite-only research credits program offers small grants to students conducting projects outside formal classes and asks recipients to acknowledge Together AI in the work.
What is Together AI's refund policy?
Together's terms say fees are non-refundable by default, and payment obligations are non-cancelable and non-proratable unless an order form says otherwise. GPU reservations are non-refundable. Canceling a fine-tuning job charges completed steps and refunds only the uncompleted portion.
Do Together AI credits expire?
Purchased prepaid credits do not currently expire. That is separate from refundability: credits can remain usable without being refundable, and credits bought after an invoice cannot clear an older past-due balance.
Did Together AI pricing change in 2026?
Yes. Together AI launched Provisioned Throughput on July 8, 2026, adding reserved token capacity at $0.05 per PTU-minute, a one-month minimum, and a 99% uptime SLA. The current serverless catalog and dedicated hardware rates have also moved since June, which is why this page carries an August 22 verification date.
Is Provisioned Throughput cheaper than Together serverless?
Not automatically. At public list rates, one fully occupied PTU maps almost exactly to the same 30-day value as its Together serverless capacity. Underuse raises the effective token price. PTUs buy reserved capacity and an SLA; a negotiated commitment discount can improve the rate, but Together does not publish that schedule.
Get the AI Tools Map for Business Owners
The AI Tools Map for Business Owners turns model prices into a practical adoption stack, with the quality floor, routing role, and cost trigger for each tool. Subscribe to get the next edition free.
Aug 22, 2026







