Gemini 3 review (2026): the benchmarks, the price, and which model to actually use

Gemini 3 reviewed with current benchmarks and prices: what each plan costs, where it beats Claude and GPT, and which model in the line to use.

Monday, June 22, 2026Omid Saffari
Gemini 3 review (2026): the benchmarks, the price, and which model to actually use

Google's Gemini 3 is now a fast-moving model family: 3.8 Flash is the stable workhorse, 3.1 Pro remains the preview reasoning model, and two 3.8 Live endpoints handle real-time voice. The honest review is not "is Gemini 3 good." It is which endpoint fits the job, what it costs at current rates, and whether Google's speed-to-price tradeoff is worth the migration.

The verdict in one read

Gemini 3 is a real frontier model line, and the developer free tier is capable enough to evaluate before you pay. The catch is that "Gemini 3" is a family with very different members, and the endpoint you land on by default is rarely the one you want.

If you live in Google's world already (Gmail, Docs, Search, Android), Gemini 3 is easy to adopt. In the United States, Google AI Plus is now $9.99/month, while the Gemini 3.8 Flash API is $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. That introductory API price, not the consumer subscription, is the sharpest value argument.

Where it is weaker: raw consistency on long, high-stakes tasks, and the kind of "just trust its judgment" reliability that keeps people on Claude. And the lineup moves fast enough that what you tested last month is already not the current model. Pick by job, not by brand, and you do very well here.

What "Gemini 3" actually is now

"Gemini 3" is not a model, it is a generation. Four endpoints matter most for the adoption decision right now, and they behave nothing alike.

A model here is just the engine doing the thinking. "Pro" is the deep-reasoning branch, "Flash" is the faster workhorse branch, and "Live" handles real-time audio. Here is the line as verified on September 16, 2026:

ModelWhat it isStateUse it for
Gemini 3.1 ProDeep-reasoning Pro modelPreviewHard reasoning and large multimodal prompts
Gemini 3.8 FlashStable general-purpose workhorseStableCoding, agents, and high-volume API work
Gemini 3.8 LiveLow-latency audio-to-audio modelStableDirect voice tasks and fast tools
Gemini 3.8 Live Extended ThinkingBackground-reasoning voice modelStableMulti-step voice tasks and slower tools

Since September 15, 2026, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking have been generally available. Standard Live defaults to asynchronous function calls, while Extended Thinking can reason and speak progress updates as tools run. In Extended sessions, turnComplete: true ends an utterance, not the whole interaction; wait for interaction_status: "IDLE" before treating the work as finished. The Gemini 3.8 Live deployment breakdown covers the call-cost and completion-state details.

The old Gemini 3 Pro Preview endpoint shut down on March 9, 2026, and Google's current model catalog labels Gemini 3.5 Flash a legacy model. For new general-purpose builds, start with 3.8 Flash rather than either one. Google describes 3.8 Flash as its most intelligent Flash model for long-horizon software engineering, autonomous agents, and complex enterprise workflows (model page).

The family no longer shares one context limit. Gemini 3.8 Flash accepts text, images, video, audio, and PDFs with a 1,048,576-token input limit and 65,536-token output limit. Gemini 3.8 Live accepts text, images, audio, and video, outputs text and audio, and has a 131,072-token input limit plus the same 65,536-token output limit (Live model page).

Gemini app interface
Gemini

If you only remember one thing: when someone says "Gemini 3," ask which endpoint. The answer changes the medium, latency, billing, and completion signal.

The benchmarks, read honestly

Gemini 3.8 Flash is now the everyday API recommendation, but its gains come with a cost caveat. Google's September launch post reports 54.9% on HLE-Verified and says the model may take extra reasoning steps and use more tokens at higher effort. Use a lower effort setting, or keep 3.7 Flash, when compute efficiency matters more than the strongest result.

A benchmark is just a standardized test: same questions for every model, scored the same way, so you can compare. For the Pro branch, Google's February 2026 model card published this head-to-head for Gemini 3.1 Pro (Thinking, High). It is a dated comparison set, not a claim about the latest rival models in September (source):

Benchmark (what it measures)Gemini 3.1 ProClaude Opus 4.6GPT-5.2
BrowseComp (agentic web search)85.9%84.0%65.8%
LiveCodeBench Pro (competitive coding, Elo)2887n/a2393
APEX-Agents (long-horizon pro tasks)33.5%29.8%23.0%
Terminal-Bench 2.0 (computer use via terminal)68.5%65.4%54.0%
SciCode (research coding)59%52%52%
SWE-Bench Verified (real coding agent)80.6%80.8%80.0%
τ2-bench retail (tool use)90.8%91.9%82.0%
MRCR v2, 128k (long-context recall)84.9%84.0%83.8%

Read that February set like this. Where the task is find, browse, and reason over a lot of moving information, Gemini 3.1 Pro is out in front, sometimes by a wide margin (BrowseComp at 85.9% versus GPT-5.2's 65.8% is not close). On head-down software engineering, the three are effectively tied, and Claude Opus 4.6 edges it on SWE-Bench Verified. On disciplined tool use, Opus leads that specific retail benchmark by a narrow margin.

The flagship reasoning numbers back the same story. At launch, Gemini 3 Pro topped the LMArena leaderboard at 1501 Elo, scored 91.9% on GPQA Diamond (graduate-level science questions) and 37.5% on Humanity's Last Exam with no tools, a deliberately brutal expert exam where 37.5% is a strong score (launch post). Turn on Deep Think, the enhanced reasoning mode, and those climb to 41.0% on Humanity's Last Exam and a standout 45.1% on ARC-AGI-2, a test built specifically to resist memorization and reward genuine novel problem solving.

What it costs, with the math

Gemini 3's sharpest price advantage is now the temporary 3.8 Flash API rate, not the consumer subscription. I rechecked Google's live model, pricing, plan, and release pages on September 16, 2026.

On the consumer side, Google's current U.S. plans page shows:

  • No Google AI plan: baseline Gemini app limits and 15 GB of Google account storage.
  • Google AI Plus, $9.99/month: 2 TB of storage, Gemini app limits 2x higher than the baseline, and access to the Flash Thinking model.
  • Google AI Pro, $19.99/month: 5 TB of storage, Gemini app limits 4x higher than the baseline, plus the Pro model and Deep Research.
  • Google AI Ultra: higher Gemini 3.1 Pro limits and access to Deep Think.

The previous AI Plus price and storage allowance are stale. The current U.S. page shows $9.99/month with 2 TB. That can still be a useful bundle if you need the storage and Google app integration, but the decision is no longer built around an unusually low sticker price.

For developers, the API is where the value gets stark (pricing, per 1M tokens):

  • Gemini 3.1 Pro: $2.00 in / $12.00 out for prompts under 200k tokens, rising to $4.00 / $18.00 above that.
  • Gemini 3.8 Flash through December 31, 2026: $0.75 in / $3.75 out. Starting January 1, 2027, those rates become $1.50 in / $7.50 out.
  • Gemini 3.8 Live and Extended Thinking: $0.75 text input / $4.50 text output; audio is $0.005 per input minute and $0.018 per output minute.
  • Grounding with Google Search: 5,000 free requests a month shared across Gemini 3.x models, then $14 per 1,000 requests.

Normalize the comparison to the pricing unit. One million input plus one million output tokens costs $4.50 on 3.8 Flash through December 31, 2026, then $9.00 starting January 1, 2027. The introductory total is half the scheduled 2027 rate. That deadline belongs in any budget forecast; a workload that looks cheap in a September pilot doubles on New Year's Day unless Google extends the offer.

Where Gemini 3 actually wins, and where it doesn't

Pick Gemini 3 for multimodal breadth, agent workflows, live voice, and current API economics. Pass on blind adoption: the Pro endpoint is still preview, the Live API guide is still preview even though the 3.8 Live model IDs are GA, and the introductory Flash price expires.

The upside
What it does well
4 points

  • Strong introductory API economics. Gemini 3.8 Flash has a free developer tier, and paid Standard pricing is $0.75 in / $3.75 out per million tokens through the end of 2026.
  • Long context and multimodal input. 3.8 Flash accepts text, images, video, audio, and PDFs in a 1,048,576-token input window.
  • Agentic search and research. Gemini 3.1 Pro led BrowseComp in Google's February comparison set, and 3.8 Flash is explicitly built for autonomous agents.
  • A real voice split. 3.8 Live handles immediate dialogue; Extended Thinking handles multi-step work and keeps the caller informed while tools run.

The honest cons. Consistency is the recurring complaint: it can be brilliant and then oddly flat on a similar prompt, which is what some forum reviewers flagged. Google also says 3.8 Flash may spend more tokens at higher effort, so the low list price does not guarantee the lowest cost per finished task. And the lineup churns: 3 Pro has already shut down, 3.5 Flash is now legacy, and builders must choose separate 3.8 endpoints for text and live voice. That velocity is good for capability and rough on operational stability.

For the full head-to-head on the everyday assistant experience, the ChatGPT vs Claude vs Gemini vs Grok breakdown goes deeper, and the Claude vs ChatGPT vs Gemini comparison puts the three main assistants side by side.

Which Gemini 3 model should you use?

Match the model to the job, not the price tag.

  • Everyday use inside Google: start with the no-plan Gemini app. Move to AI Plus at $9.99 only when the higher limits, Flash Thinking access, Gmail integration, or 2 TB of storage solve a real constraint.
  • High-volume text, coding, or agent loops: start with Gemini 3.8 Flash. Its stable endpoint and introductory rates make it the default API choice, but budget against the January 2027 rates too.
  • Fast voice tasks: use Gemini 3.8 Live. It is built for immediate turn-taking and direct tools.
  • Multi-step voice tasks: use Gemini 3.8 Live Extended Thinking. Its background reasoning and progress speech justify the extra client-state work only when tools take time or the task needs planning.
  • The hardest multimodal reasoning: test Gemini 3.1 Pro Preview. It remains the Pro branch, but preview status should be part of the production decision.
  • High-stakes work: run your own acceptance set and require human review. A benchmark lead does not remove the need to check contracts, medical or legal drafts, and financial decisions.

The decision rule in one line: use 3.8 Flash for general API work, standard Live for immediate voice, Extended Thinking for multi-step voice, and 3.1 Pro only when deeper reasoning earns the preview risk and higher rate.

How to actually get it

Getting onto Gemini 3 is straightforward, but the right starting point now depends on whether you are chatting, building a text workflow, or shipping voice.

  1. Start free

    Open the Gemini app or gemini.google.com and sign in with a Google account. Start at the no-plan baseline and use it on your real work before paying for higher limits.

  2. Upgrade only if you hit a wall

    If the baseline limits or storage become a constraint, compare Google AI Plus at $9.99/month with Google AI Pro at $19.99/month. Plus includes 2 TB and 2x the baseline Gemini app limits; Pro includes 5 TB, 4x the baseline limits, the Pro model, and Deep Research.

  3. Get an API key for building

    Go to Google AI Studio, create a key, and call 3.8 Flash first. Move to 3.1 Pro Preview only for prompts that genuinely need the Pro branch's deeper reasoning, and price the difference before routing production traffic.

  4. Choose the Live branch deliberately

    Use 3.8 Live for direct voice tasks and fast tools. Use Extended Thinking only when the workflow needs multi-step reasoning or slow tools, then track interaction_status until it becomes IDLE.

Is Gemini 3 free?

Yes. The Gemini app has a no-plan baseline, and the Gemini Developer API offers free input and output for supported models at limited rates. In the United States, paid consumer plans start with Google AI Plus at $9.99/month; Deep Think requires Google AI Ultra.

How much does Gemini 3 cost?

Consumer pricing in the United States starts at $9.99/month for AI Plus and $19.99/month for AI Pro. For the API, Gemini 3.8 Flash is $0.75 in / $3.75 out per million tokens through December 31, 2026, then $1.50 in / $7.50 out from January 1, 2027; 3.1 Pro is $2 to $4 in / $12 to $18 out depending on prompt size.

Is Gemini 3 better than ChatGPT or Claude?

Not across every task. In Google's February 2026 comparison set, 3.1 Pro led on BrowseComp, LiveCodeBench Pro, APEX-Agents, Terminal-Bench 2.0, SciCode, and long-context recall, while Claude Opus 4.6 narrowly led on SWE-Bench Verified and the retail tool-use test. Those are vendor-published benchmarks, so use your own acceptance set for the final decision.

What is Gemini 3 Deep Think?

An enhanced reasoning mode that spends more compute on hard problems. Google's launch evaluation reported 45.1% on ARC-AGI-2 with code execution. It is gated to Google AI Ultra and is overkill for everyday tasks.

Gemini 3 Flash vs Pro, which should I use?

Use stable Gemini 3.8 Flash for most API work, including coding and agent loops. Use Gemini 3.1 Pro Preview when your own evaluation shows that deeper reasoning justifies the preview endpoint and its higher token price.

Want this kind of current, no-hype breakdown of the tools that actually matter, with the real prices and the honest trade-offs, every week? Join the newsletter.

Last Updated
Sep 16, 2026
Category
AI

Prefer this site in Google

Add omidsaffari.com as a preferred source in Google Search

Mark omidsaffari.com as preferred and Google lifts it in Top Stories, AI Overviews and AI Mode for you.

Related Articles
Newsletter

One letter, every Sunday.Working systems, not hot takes.

Weekly. No spam. Unsubscribe anytime.