Google debuted its latest AI model, Gemini 3.8 Flash on Wednesday. Google said the model excels at coding and on some benchmarks, its performance equalled that of larger models from rival AI companies, but completed them at a much lower cost. The model’s release comes just three weeks after the release of its predecessor, Gemini 3.7 Flash. And, overall, the company has released no less than four Gemini Flash models since May. Flash is the designation Google uses for the smallest, and fastest versions, of the models it produces. They are generally designed for users seeking responses at a low cost, without sacrificing too much cognitive power. But Google’s flagship Gemini 3.5 Pro model, which CEO Sundar Pichai said would arrive in June, is still missing in action. That has led many AI industry insiders to question whether Google is still able to catch up to the frontier AI of technology. “I really hate to say it, but…gemini who?” Meta chief AI officer Alexander Wang trolled Google on X this week following the release of his own company’s new Muse Spark 1.3 model, which leapfrogged all of Google’s models on a closely-watched performance benchmark. On that benchmark, the Artificial Analysis Intelligence Index, Google’s best performing model, Gemini 3.8 Flash, now ranks in 10th place.
Google expected Gemini 3.5 Pro to ship in June. It was still undergoing testing in July, and Google’s website still lists it as “coming soon.” Internal candidates were discarded because they did not improve enough over Flash, the Wall Street Journal reported.
Google, however, has begun pointing to the pace at which it has been able to release new Flash models as a sign that it is increasingly in the forefront when it comes to a key AI building method that the AI industry is cares deeply about: recursive self-improvement, or RSI for short. RSI refers to an AI model that can improve and optimize its own code, spawning ever-more-capable versions of itself. For many AI researchers, RSI has been a long-term goal, while for many AI safety experts, it has been among their leading fears. That’s because they worry that once AI models can self-improve with little human input or oversight, there could be an “intelligence explosion” that rapidly leads to models that are far more intelligent than all of humanity, with dire consequences for humans. Many AI labs have begun dipping their toes into AI building techniques that are somewhat similar to RSI—using one generation of AI models to help design and build the next, but usually with a good deal of human oversight and input. Google DeepMind researcher Shunyu Yao wrote in an X post on Gemini 3.8 Flash’s debut that the model represented “one small step for model, one giant leap for RSI.”
Google said in its release announcement, Gemini 3.8 was “further accelerated” by long-running AI-agent loops that “recursively evaluate and refine the underlying models.”
That is a more direct claim about model development than Google made in its previous Flash announcements. In May, Google used a “self-improvement loop” for two agents building and playing a game. In August, it described a three-agent loop helping train a robotics model. For Gemini 3.8 Flash, Google says the loops refined the Gemini models themselves. It is unclear what Google’s use of “self-improvement loops” portends about its use of similar techniques for larger AI models. It could be that the smaller Flash versions of Gemini are easier to improve on using these methods than the larger Gemini Pro versions.
Flash models require less computing capacity to modify, according to people cited by the Journal. That allows several research teams to test different approaches in parallel. Changes to Google’s larger Pro models require more resources.
The commercial incentives also favor Flash. Pichai called Flash Google’s “workhorse” series and said it hit the “sweet spot of performance and cost.” In its Q2 earnings call, Google said its model APIs were processing about 22 billion tokens per minute, up from 16 billion one quarter earlier. The company said computing supply remained constrained.
That combination gives Google a commercial reason to keep improving Flash as API demand grows and computing capacity remains constrained. Its effort settings also let customers trade performance for speed and cost.
It could also be the case, however, that Google is hoping to use RSI-like techniques to soon jump back to the front of the AI race, creating models that would be more capable than Anthropic’s Mythos 5 or OpenAI’s Astra.
What is known is that Gemini 3.8 Flash’s release follows months of internal investment in coding and reinforcement learning. Since the beginning of the year, Google has directed more researchers’ time and computing resources toward improving Gemini’s coding abilities, The Wall Street Journal reported.
By April, Google had assembled a coding strike team, with cofounder Sergey Brin and DeepMind technology chief Koray Kavukcuoglu directly involved, The Information reported. Brin told employees that improving coding was a step toward self-improving AI and urged DeepMind to turn its models into “primary developers” of code, the publication reported.
Brin told employees that stronger coding models were a step toward AI systems capable of improving themselves. In an internal memo, he urged Google to close its gap in agent execution and turn its models into “primary developers” of code, according to the story.
Google’s faster Flash cadence has not extended to its flagship Pro line.
Google said Gemini 3.8 Flash improved over prior versions in software engineering, agentic tasks and multi-step reasoning. It also released a cybersecurity version through Fairwind, a controlled-access program for governments and national cyber authorities, critical infrastructure operators and organizations that maintain widely used software.
Independent tests place Gemini 3.8 closer to more expensive frontier models on some coding tasks. DeepSWE v1.1 tests whether coding agents can turn a short request into a working change in an existing codebase. Gemini 3.8 Flash at high effort and Anthropic’s Claude Opus 5 at maximum effort each passed about 74% of scored runs. Their reported error ranges overlapped.
Across all scored attempts, Gemini averaged $2.36 in model-use costs per task. Opus averaged $11.84. Gemini used more output tokens and took more agent steps, yet its lower token price kept its average task cost below Opus.
That pattern explains Google’s claim that Gemini 3.8 “works harder.” The model takes additional reasoning steps and calls tools repeatedly when it faces a difficult task.
More work can also make it more expensive than its predecessor. Google kept Gemini 3.8 Flash’s introductory API price at 75 cents per million input tokens and $3.75 per million output tokens. Artificial Analysis found that at high reasoning, Gemini 3.8 Flash cost about 40% more per task than Gemini 3.7 Flash, despite identical token prices. It attributed the increase to 30% more output tokens per task and more turns in agentic evaluations.
The post Google shipped four Gemini Flash models in 106 days. But its flagship frontier model is still nowhere to be seen. appeared first on Fortune.




