Google released Gemini 3.8 Flash today. It scores 59 on the Artificial Analysis Intelligence Index in its high-reasoning configuration — up three points on Gemini 3.7 Flash, and well above the 36 median for comparable models. The gain comes mostly from agentic evaluations: tool use, coding, multi-step real-world tasks. Google says it beats Anthropic’s Opus 5 and OpenAI’s GPT-5.6 Sol on a number of benchmarks while costing a fraction as much.
Pricing is $0.75 per million input tokens and $3.75 per million output — against category medians of $1.75 and $10.00. That is the identical introductory rate Google put on Gemini 3.7 Flash. Introductory is the operative word: the posted rate rises to $1.50 and $7.50 from January 2027.
Now the calendar. Gemini 3.7 Flash landed roughly three weeks ago. 3.8 Flash is the fourth Flash model in under four months. No frontier Gemini release has accompanied any of them.
Three weeks is not a product cycle. It is a deployment problem.
If you are a consumer, faster and cheaper every three weeks is unambiguously good. If you run anything in production, a three-week cadence is a tax, and it is a tax nobody budgets for.
Every release forces the same decision: stay pinned to a model that is now two versions old and getting cheaper competitors every month, or re-run your evaluations, re-tune your prompts, re-check your latency budget and migrate. Do that every three weeks and you have hired someone to do nothing but chase model versions. Skip it four times and you are a year behind on capability and paying above-market rates.
Our take: The cadence is the strategy. Google is not trying to win a benchmark argument; it is trying to make the cheap tier a moving target that competitors have to re-price against every few weeks, while the expensive frontier tier stays quiet. For buyers, that means the version number in your config file has become the least stable part of your stack — and the only defence is an evaluation suite you can re-run in an afternoon. If migrating models is a multi-week project for you, this release cadence will eat you. If it is a Tuesday, it is free money.
Read the cost line properly
Cheaper per token is not cheaper per job. Artificial Analysis puts Gemini 3.8 Flash on the intelligence-versus-cost frontier at about $0.58 per task — comparable to GPT-5.6 Terra at roughly $0.53, but around 40% higher than its own predecessor. The reason is that average output climbed roughly 30% to about 48,000 tokens per task. The model thinks more. Thinking is billed.
The other number worth checking against your own requirements: time to first token is about 13.4 seconds, against a roughly 3-second median for reasoning models in the same price band. Throughput afterwards is fast, at around 305 tokens per second. For batch work that is irrelevant. For anything a human is waiting on, it is the whole experience.
What to watch
- January 2027. Introductory pricing doubles. Anyone modelling unit costs on $0.75 is modelling a promotion.
- The missing frontier model. Four budget releases and no flagship is a conspicuous pattern.
- Whether 3.7 Flash gets a retirement date. Shipping fast is one thing; switching models off is what actually breaks applications — as AWS customers are being reminded this month.
- Tokens per task, not price per token. The reasoning-token creep is where the cost quietly moved.
