in

Google Launches Gemini 3.8 Flash With a Price Rise Set for January

Google released Gemini 3.8 Flash on 2 September 2026, its third Flash model in six weeks, and kept the introductory price of the version it replaces: $0.75 per million input tokens and $3.75 per million output tokens. A footnote at the bottom of the announcement sets an end date on that. The introductory price expires on 31 December 2026, and from 1 January 2027 the rate becomes $1.50 per million input tokens and $7.50 per million output tokens, exactly double.

Google published the release alongside a second model, Gemini 3.8 Flash Cyber, which is not available to the general public at all.

Hosting 75% off

Three Flash releases in six weeks

The pace is the first thing worth noticing. Gemini 3.7 Flash shipped three weeks before this one. Google describes 3.8 Flash as its most intelligent workhorse model and says it often approaches the performance of higher cost frontier models, particularly on long running software engineering work measured by the DeepSWE v1.1 benchmark. The model scores 54.9% on HLE-Verified, a test of multi step reasoning across science, humanities and professional subjects, and Google reports gains on finance and legal agent benchmarks run by Vals AI.

It is available now through the Gemini API in Google AI Studio and Android Studio, in Google Antigravity, and in Gemini Enterprise. Consumers on Google AI Pro and Ultra plans get it inside the Gemini app, AI Mode in Search, and Gemini in Google Sheets.

The model works harder, and that shows up on the bill

Buried in Google’s own write up is a caveat that matters more than the headline price. The performance gains come from what Google calls a core design choice: on complex tasks the model executes extra reasoning steps and calls tools repeatedly. In Google’s words, it “might use more tokens to maximize performance, especially at higher effort levels.”

Read that next to the price and the arithmetic gets less friendly. A model that costs the same per token but spends more tokens on the same job is not cheaper to run. Google’s own advice is to drop to lower effort levels when compute cost is the binding constraint, or to stay on Gemini 3.7 Flash, which remains fully supported for efficiency first workloads.

For anyone running an AI product or client service on thin margins, that is the practical decision this week. Benchmark the new model against your actual workload rather than the announcement, count total tokens consumed and not price per token, and note the January date in whatever spreadsheet holds your cost assumptions. We covered a similar pattern when OpenAI and Anthropic cut rates while DeepSeek raised them, and the lesson held then too: advertised price and delivered cost are different numbers.

Who gets the cybersecurity version?

Gemini 3.8 Flash Cyber shares the same underlying model but ships with looser safety restrictions around cybersecurity work, which is why Google is gating it. Access runs through a new Fairwind Program aimed at trusted government authorities, critical infrastructure operators and software maintainers, who have to apply.

The capability numbers explain the caution. On CWE-Bench, an external benchmark for patching, the model scores 47.2% pass@1 against a leading frontier model’s 47.8%, at a much lower cost. On an internal Google benchmark covering vulnerabilities across 20 programming languages it exceeds a 70% success rate. Google’s Chrome security team reports it produced 2.6 times more correct patches than the best commercial models, which are considerably larger. Google’s cloud vulnerability research team says it used the model to find a critical foundational flaw in under two hours, work that normally takes months.

A system that finds and fixes flaws that fast can also find them for other reasons. Google says it deliberately prioritised patching over exploitation during development, and that the public 3.8 Flash ships with safeguards against cyber offence and against chemical, biological, radiological and nuclear misuse.

Where that leaves builders outside the program

Two tiers are forming, and the split is not about money. The public model is fast, cheap for now, and deliberately limited on offensive security. The gated model is the capable one, and reaching it requires being a government, an infrastructure operator or a recognised maintainer. A small agency in Karachi or a solo developer in Dubai sits firmly in the first tier no matter what they are willing to pay.

That is a defensible call on Google’s part, and OpenAI and Anthropic have made comparable moves on cyber capability in the same window. It also means the security tooling available to small operators will keep lagging the tooling available to large institutions, while the attackers those tools defend against are drawing on whatever they can obtain. The gap is worth watching as these access programs expand.

Sources: Google, “Introducing Gemini 3.8 Flash and 3.8 Flash Cyber” (company announcement); Google, Fairwind Program; The Register.

Hosting 75% off

Written by Ahmed Shaami

Nvidia Confirms $12.93 Billion Deal to Buy Hugging Face