MiniMax AI (August 2026 Review)
MiniMax AI review: M3, H3, Code, Hub, current pricing, US open-weight limits, and who should choose ChatGPT, Claude, or Runway instead.
- MMiniMax

MiniMax AI is worth considering when access to M3 coding, H3 video, Speech 2.8, Music 3.0, and low-cost APIs through one vendor is more valuable than one polished chat app. Plans start at $20 a month, H3 costs $0.08 per second at 768P, and the biggest catch for US teams is that H3's open weights require separate authorization even though its hosted API is global.
What MiniMax AI Actually Is
MiniMax is an AI model company with several products that share models and billing but do not behave like one unified assistant. M3 is its long-context language and coding model. MiniMax Code is the agent interface for coding and general work. Hub is a desktop creative canvas that coordinates text, image, video, and audio agents. H3 is the new video model. Speech 2.8, Music 3.0, image-01, and the Open Platform APIs cover the rest of the media stack. That breadth is the appeal, but it also explains the confusing first encounter: the homepage, Code, Hub, Audio, Hailuo, Token Plan, and API console are different doors into the same company. A buyer should review MiniMax as an ecosystem, not as a single ChatGPT replacement.

MiniMax AI vs the Alternatives at a Glance
MiniMax wins on breadth per dollar, while ChatGPT, Claude, and Runway each offer a cleaner answer to one narrower buying problem. Prices below were verified on the vendors' live pages on August 6, 2026.
The choice flips on the center of gravity. Pick MiniMax when a technical workflow uses two or more of its strengths, such as M3 plus H3 or H3 plus Speech 2.8. Pick the specialist when one job dominates. Paying for MiniMax's breadth to use only chat creates complexity without a return.
Who MiniMax Is For and Who Should Skip It
MiniMax fits technical builders and creative operators who can turn low unit prices into repeatable workflows. It is a weaker default for buyers who judge a tool mainly by the smoothness of one consumer app.
Choose MiniMax for high-volume model and media work
A solo technical builder can use M3 through MiniMax Code or an existing coding client, then call H3, speech, music, and image APIs from the same platform. That combination matters when a product needs generated help-center narration, short video assets, image variants, and an agent that maintains the pipeline. The value comes from fewer vendors and low metered rates, not from one magical interface.
A funded founder or mid-market CTO should look at MiniMax when long repository context and cost control matter together. M3's hosted context reaches 1M tokens, the Token Plan supports several coding tools, and pay-as-you-go starts at $0.30 per million input tokens below 512K context. The right trial is one expensive, repeatable route with an acceptance rubric, not a company-wide model swap.
A creative operator should shortlist MiniMax when native-audio video is central. H3 produces 4 to 15 second clips with stereo audio and can take text, images, video, and audio as references. Hub packages those capabilities into a canvas. That is a more distinctive reason to choose MiniMax than generic chat.
Choose ChatGPT for one general-purpose workspace
ChatGPT is the better default for a nontechnical operator who wants files, search, voice, projects, scheduled tasks, and a broad assistant in one familiar place. The live individual plans start with Go at $8 monthly, while Plus costs the same $20 monthly as MiniMax Plus. MiniMax can cover more model modalities, but the buyer must understand which surface and balance pays for each job.
Choose Claude when coding judgment matters more than unit price
Claude is the safer shortlist for a senior builder whose cost is dominated by review time and failed changes rather than tokens. Claude Pro costs $20 month-to-month or an effective $17 monthly when $200 is paid annually. MiniMax M3 deserves a measured trial for long-context and high-volume work, but vendor-run benchmarks are not a substitute for comparing accepted patches, retries, tool errors, and reviewer minutes on the same repository.
Choose Runway when the deliverable lives in a video editor
Runway is the clearer fit when a creative team wants a video-first product, model choice, shared production economics, and an established credit workflow. Standard costs $15 monthly or $12 monthly with annual billing and includes 625 monthly credits, which Runway equates to 52 seconds of Gen-4.5. MiniMax H3 can be less expensive per generated second, but raw generation cost does not buy a mature review and editing process. The broader AI video generator comparison is the better map when video is the only buying question.
The Four MiniMax Capabilities That Matter
MiniMax earns a place on the shortlist through four concrete workflows: long-context agent work, reference-driven video, coordinated creative production, and low-cost voice generation.
1. MiniMax M3 and Code make long-context agent work inexpensive
MiniMax M3 is a coding and agent model with a hosted context window up to 1M tokens and a guaranteed minimum of 512K. MiniMax says it accepts native multimodal input and is built for task decomposition, tool calls, and extended runs. Those claims become useful when the job carries a large repository, issue history, screenshots, and tool output through one session.

The strongest evidence on the page is still vendor-run. MiniMax reports that M3 reproduced an academic paper over nearly 12 hours, producing 18 commits and 23 experimental figures. It also reports a roughly 24-hour CUDA optimization run with 147 submissions and 1,959 tool calls. Those examples show the intended workload. They do not prove M3 will beat Claude or ChatGPT on your codebase.
MiniMax Code is the ready-made surface for that model. It adds persistent memory, reusable skills, Agent Teams, and desktop or web access. A Token Plan can run Code directly, or provide a separate sk-cp key for a listed OpenAI-compatible coding tool such as Cursor, Codex CLI, Cline, Claude Code, or OpenCode.

Here is the workflow that makes the economics measurable:
Choose one repository task
Use a bounded issue with existing tests, a known failure, and a reviewer who can accept or reject the result. A repository-wide migration is too broad for the first comparison.
Keep the setup constant
Connect M3 to the same coding client used by the incumbent model. Preserve tool permissions, test commands, timeouts, and the acceptance rubric so the comparison measures the model rather than a new interface.
Price the complete request
A request with 200,000 uncached input tokens and 20,000 output tokens costs $0.084 at standard M3 rates below 512K: $0.06 for input plus $0.024 for output. A 700,000-input and 30,000-output request costs $0.492 at the higher long-context rate.
Count accepted outcomes
Record model spend, elapsed time, retries, failed tool calls, changed files, and reviewer minutes. Divide total cost by accepted tasks. Cheap output that needs repair is not cheap work.

2. MiniMax H3 combines references, video, and native stereo audio
MiniMax H3 is the standout reason to look beyond the chat and coding market. It accepts a multimodal brief and generates a 4 to 15 second video at 24 FPS with 32 kHz stereo audio. The reference mode allows up to 9 images, 3 videos, 3 audio clips, and 12 files total, which is enough to specify a subject, visual language, motion reference, and voice reference in one job.

A product-launch shot can be structured without pretending that a single text prompt will carry the whole brief:
- Provide the product stills as reference images and a short movement reference as video.
- Add the narration or voice character as an audio reference. Audio cannot be the only reference input.
- Choose a 4 to 15 second duration and a supported ratio. Text-to-video requires an explicit ratio; image-to-video follows the input image.
- Generate at 768P for the selection pass, then move the accepted shot to 2K.
The meter is direct. A 10-second 768P output costs $0.80. The same duration at 2K costs $1.30, and a 15-second 2K output costs $1.95. A 10-second 2K generation that also consumes a 10-second reference video costs $2.60 because MiniMax bills both output duration and reference-video duration at $0.13 per second. Audio references are free, the first 5 image references are free, and later images cost $0.04 each.
ComfyUI shipped native H3 templates for text-to-video, image-to-video, and reference-to-video at launch. That makes the 768P Base checkpoints usable in a node workflow, but it does not make the full hosted system local. The distinction becomes the central limitation later in this review.
3. MiniMax Hub coordinates a creative brief across media
MiniMax Hub is the attempt to turn the company's separate models into one production canvas. Its public workflow starts with a brief, breaks the work into script, storyboard, image, video, music, voiceover, and editing tasks, and can launch copy, image, video, and audio agents in parallel.

The practical workflow is a small campaign, not a one-shot masterpiece. Give Hub a product brief, target audience, approved claims, reference assets, and deliverable list. Let the copy agent draft the script and captions, the image agent establish visual candidates, the video agent build short sequences, and the audio agent prepare voiceover or music. Review each stage before the next model spends more credits. Export selected assets to the professional tool that will finish them.
That last review step matters. Agent orchestration reduces tool switching, but it does not remove art direction, rights clearance, brand checks, or edit decisions. Hub is most valuable when it turns a repeated campaign recipe into a reusable skill. It is less compelling when the team already has a mature editor-centered workflow and only needs one generation model.
4. MiniMax Speech 2.8 makes narration costs easy to predict
MiniMax Speech 2.8 supports native sound tags and voice cloning from a 10-second sample according to the current release page. The synchronous API accepts up to 10,000 characters per request, while the asynchronous route accepts up to 1M characters. MP3, PCM, FLAC, and WAV are supported, with WAV limited to non-streaming use.

For a 50,000-character training narration, speech-2.8-turbo costs $3 and speech-2.8-hd costs $5. A rapid voice clone adds $1.50 on first use, while a designed voice costs $3. The workflow is straightforward: secure a properly authorized voice sample, generate a short approval clip, lock pronunciation and pace, then send the full script through the asynchronous endpoint and retain the output before its temporary delivery link expires.
Music and image rates make small supporting batches similarly legible. Twenty Music-3.0 songs cost $3 at $0.15 each, and 100 image-01 outputs cost $0.35 at $0.0035 each. Those rates make exploration cheap. Selection, rights, consistency, and finishing remain the expensive parts of the outcome.
MiniMax Pricing: Every Current Tier
MiniMax pricing is inexpensive at the model layer and complicated at the account layer. The public system has Token Plans, separate prepaid credits, pay-as-you-go API balance, Audio Subscriptions, and legacy Video Packages. Every rate below was checked against the live MiniMax pricing overview and linked pages on August 6, 2026.

Token Plan subscriptions
The prepaid Token Plan credit packs are Starter at $5 for 5,000 credits, Growth at $25 for 25,000, and High Volume at $100 for 100,000. Purchased credits remain valid for one year. They are not the same thing as the standard pay-as-you-go API balance.
The quota language deserves as much attention as the large token estimates. Included usage is controlled by rolling 5-hour and weekly windows, unused included quota does not carry into the next billing cycle, and peak-hour capacity can tighten. MiniMax explicitly recommends pay-as-you-go for production.
M3 pay-as-you-go API
MiniMax pay-as-you-go separates Standard and Priority service, with a price step when input exceeds 512K tokens.

At Standard priority and no more than 512K input, M3 costs $0.30 per million input tokens, $1.20 per million output tokens, and $0.06 per million prompt-cache reads. Above 512K, those rates become $0.60 input, $2.40 output, and $0.12 cache reads.
Priority is 1.5 times Standard. At or below 512K, it costs $0.45 input, $1.80 output, and $0.09 cache reads per million. Above 512K, it costs $0.90 input, $3.60 output, and $0.18 cache reads.
M2.7 remains $0.30 input, $1.20 output, $0.06 cache read, and $0.375 cache write per million tokens. M2.7-highspeed doubles input and output to $0.60 and $2.40 while retaining the same cache rates.
H3 and media API rates
MiniMax H3 costs $0.08 per output second at 768P and $0.13 per output second at 2K. Regenerating an existing 768P result to 2K costs $0.05 per output second, with the original input materials billed again. The first 5 reference images are free; later images cost $0.04 each during generation and $0.025 each during regeneration. Reference audio is free. Reference video is metered per second at the selected output resolution's rate.
Speech-2.8-turbo costs $60 per million characters, and speech-2.8-hd costs $100. Rapid voice cloning is $1.50 per voice and Voice Design is $3 per voice. Music-3.0 is $0.15 for up to 5 minutes, with a free 3-RPM model. Lyrics generation or editing costs $0.01 per song. Image-01 is $0.0035 per image, API-vlm is $0.01 per request, and server-side web search is $0.01 per request.
Every Audio Subscription tier
MiniMax sells five fixed speech subscriptions plus custom pricing:
- Starter: $5 monthly, $13.50 for 3 months, or $48 yearly; 100,000 audio points/month, 10 RPM, and 10 voice slots.
- Standard: $30 monthly, $81 for 3 months, or $288 yearly; 300,000 points, 50 RPM, and 100 voice slots.
- Pro: $99 monthly, $267 for 3 months, or $950 yearly; 1.1M points, 200 RPM, and 250 voice slots.
- Scale: $249 monthly, $672 for 3 months, or $2,390 yearly; 3.3M points, 500 RPM, and 500 voice slots.
- Business: $999 monthly, $2,697 for 3 months, or $9,590 yearly; 20M points, 800 RPM, and 800 voice slots.
- Custom: negotiated capacity, unlimited RPM/TPM, priority updates, and negotiated service guarantees.
The subscription page names points rather than promising a universal character conversion. Compare the intended workload in the billing console before treating the point total as a fixed number of narration minutes.
Every legacy Video Package tier
MiniMax's one-month Video Packages are Standard at $1,000 for 3,760 video points and 20 RPM, Pro at $2,500 for 9,920 points and 30 RPM, Scale at $4,500 for 18,900 points and 40 RPM, and Business at $6,000 for 26,780 points and 50 RPM. Custom pricing offers negotiated points and unlimited RPM/TPM.
These packages support the Hailuo series and do not support H3 yet. Unused package quantity resets to zero when the pack expires. A buyer drawn in by H3 should not assume that a large Hailuo pack discounts H3 usage.
The Limitations That Change the Decision
MiniMax has four limitations serious enough to change who should buy it and how it should be deployed.
H3's open weights are partial and territorially restricted
MiniMax H3 is open-weight only at the Base layer. The released Base checkpoints produce 768P. H3-Context-IR, which MiniMax describes as critical to final quality, remains a hosted service. H3-Regenerate-2K is also not open, and the initial release does not include the model's sparse-attention implementation. MiniMax's own SGLang commands specify 4 GPUs for each task checkpoint.
The license is an even larger constraint for the site's US-first audience. The standard H3 Community License defines the United States, European Union, United Kingdom, and South Korea as excluded territories. Organizations in those regions can apply for formal authorization, and the hosted API remains globally available. Commercial products generating more than $20M in annual revenue also require prior written authorization.
That means "downloadable" is not the same as "ready for a US production deployment." Use the hosted API for evaluation, and get the correct authorization before building a local product around the weights. The wider open-weight model guide explains why files, license, hardware, and a reproducible serving path are four separate gates.

Token Plan is not a production SLA
MiniMax says Token Plan is designed for individual, interactive developer use and recommends pay-as-you-go for production. Rolling 5-hour and weekly limits, non-carrying quota, and dynamic peak throttling are acceptable for a personal coding agent. They are not a capacity commitment for a customer-facing service.
Use Token Plan to learn the model and support interactive work. Move a production route to pay-as-you-go with explicit timeouts, retries, observability, and spend caps.
Billing is fragmented across products
MiniMax's low rates sit behind several ledgers. Token Plan quota, Token Plan credits, pay-as-you-go balance, audio points, and video points are not one interchangeable pool. Video Packages exclude H3. Team seat pricing is not public. A finance owner cannot estimate the monthly total from the Plus, Max, and Ultra cards alone.
Build the forecast by workload: M3 input/output, H3 output and reference duration, speech characters, music songs, image count, and any specialist subscription. Then assign each line to its actual balance. This is tedious, but it prevents a cheap headline rate from hiding stranded credits or a second bill.
The strongest performance evidence is vendor-reported
MiniMax publishes ambitious long-horizon examples and benchmark claims for M3. They are useful for choosing a test, not for skipping one. The live M3 page also uses "open-weight" language while its access section says full open sourcing is coming soon. Treat the exact files, license, serving framework, and release date as items to verify again before planning private M3 infrastructure.
The Kling AI review applies the same standard to video: the model earns a place through accepted shots and controllable workflows, not its best launch demo. MiniMax deserves the same discipline.
- M3, H3, speech, music, image, and agent surfaces under one vendor
- Low M3 API rates and transparent per-second H3 pricing
- H3 reference inputs combine images, video, and audio with native stereo output
- MiniMax Code works directly or through several established coding clients
- Live pricing exposes enough detail to calculate workload-level cost
- H3's local release omits Context-IR, 2K regeneration, and sparse attention
- Standard H3 open-weight rights exclude the US, EU, UK, and South Korea without authorization
- Token Plan quotas can tighten and are not intended for production
- Separate balances and packages make total cost harder to forecast
- The public Team price is not disclosed, and legacy video packs exclude H3
MiniMax AI Verdict
MiniMax is a strong value choice for a technical buyer with a specific multimodal workload, not the default AI subscription for everyone.
Choose MiniMax when at least two of these are true:
- M3 handles a repeated repository or agent route cheaply enough to clear the acceptance bar.
- H3's reference-driven video and native audio remove a separate generation step.
- Speech, music, or image API volume makes the platform's micro-rates meaningful.
- Your team can manage separate product surfaces, balances, and deployment rules.
Skip MiniMax when one polished assistant is the whole requirement, when reviewer time makes model judgment more important than token price, or when a video team needs an editor-centered production system. ChatGPT, Claude, and Runway win those three cases respectively.
The cost rule is explicit. Start M3 on pay-as-you-go. For the 200K-input and 20K-output example, Plus becomes cheaper around 239 requests per month before cache savings. Keep customer-facing traffic on pay-as-you-go because MiniMax itself does not position Token Plan as production capacity. Start H3 at 768P, pay for 2K only after a shot survives review, and include reference-video duration in the cost estimate.
The deployment rule is equally explicit. A US, EU, UK, or South Korean organization should use the H3 API or secure formal open-weight authorization before local deployment. Do not buy hardware from the word "open."
MiniMax AI FAQ
How much does MiniMax AI cost?
MiniMax Token Plans cost $20/month for Plus, $50 for Max, and $120 for Ultra, with annual prices of $200, $500, and $1,200. M3 pay-as-you-go starts at $0.30 per million input tokens and $1.20 per million output tokens below 512K context, while H3 starts at $0.08 per output second at 768P.
Is MiniMax AI better than ChatGPT?
MiniMax is better for a technical buyer combining low-cost M3, H3 video, speech, music, and APIs. ChatGPT is better when the goal is one polished general workspace for files, search, voice, projects, and everyday knowledge work.
Is MiniMax AI safe to use?
MiniMax H3 applies automated moderation to submitted text, images, and video, but its model card says false positives and false negatives remain possible. Keep sensitive material out of an evaluation until the relevant privacy, security, and contract terms meet your requirements, and enforce permissions outside the model.
Is MiniMax worth it?
MiniMax is worth it when a measured M3, H3, speech, or mixed-media route crosses the cost threshold and still meets the acceptance bar. It is not worth adding several MiniMax surfaces for occasional chat that a simpler subscription already handles.
Can MiniMax H3 run locally?
The 768P H3 Base checkpoints can run locally, and ComfyUI provides native workflows. The official MiniMax SGLang example uses 4 GPUs, while the quality-critical Context-IR and 2K regeneration stages remain hosted. The standard open-weight license also excludes the US, EU, UK, and South Korea unless formal authorization is granted.
Want the current model picks mapped to the work they fit? Get the AI tools map for business owners, free.
Get the AI tools map for business owners
A plain-English map of which AI models and tools fit which job, updated as prices and capabilities change. Free to subscribers.
Aug 6, 2026







