Modelspublished

Google May Ship Gemini 3.8 Flash 20 Days After 3.7, With Coding Edge Unproven

Google engineers reportedly preferred the internal model to an Anthropic Opus system in Jetski. Without the test design or public model materials, that result is a signal of internal usefulness rather than a broad coding verdict.

By 3 min read
Google May Ship Gemini 3.8 Flash 20 Days After 3.7, With Coding Edge Unproven
Google May Ship Gemini 3.8 Flash 20 Days After 3.7, With Coding Edge Unproven

Listen to this story

The audio brief

About 1:36
0:001:36
Read transcript
Google may release Gemini 3.8 Flash on September 2, just 20 days after Gemini 3.7 Flash. That would be one of the fastest numbered updates yet for Google’s lower-cost model line. The Wall Street Journal also reports that Google employees preferred the new system—known internally as Skimaki—to an unspecified Anthropic Opus model while using Google’s Jetski coding tool. That sounds important for developers, but it is not a public coding verdict. The test did not disclose its prompts, repositories, evaluators, results, methodology, or even the exact Opus version. Preference inside one company can reflect a model’s speed, interface, instruction style, or integration—not necessarily more correct, secure, or maintainable code. There is a useful public baseline. Gemini 3.7 Flash scored 43.6 percent on Google’s FrontierCode 1.1 evaluation, versus 42.7 percent for Claude Sonnet 5. But Google’s own comparisons showed it trailing OpenAI’s GPT-5.6 Terra on DeepSWE and Terminal-bench software-engineering tests. The economics matter too: Gemini 3.7 Flash launched at 75 cents per million input tokens and three dollars 75 per million output tokens, through the end of 2026. So the near-term signal is rapid iteration, not proven superiority. The key question is whether Gemini 3.8 arrives with public benchmarks, safety evaluations, pricing, and a model card that let developers judge the claim across real coding workloads.

Story brief

3 key points

The potential September 2 launch would make Gemini 3.8 Flash one of Google’s fastest follow-ons, arriving just 20 days after 3.7. Its reported coding advantage is based only on Google employees preferring it to an unidentified Anthropic Opus model inside Google’s Jetski tool—not public benchmarks. Developers therefore have no confirmed pricing, safety data, or reproducible comparison yet. Gemini 3.7’s $0.75 input...

  1. 01

    Gemini 3.7 Flash launched August 13; a September 2 successor would establish a 20-day numbered-release interval.

  2. 02

    Google’s introductory Gemini 3.7 pricing is $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026.

  3. 03

    The reported Jetski test did not disclose prompts, evaluators, repository tasks, Opus version, results, or methodology.

Google could release Gemini 3.8 Flash as soon as September 2, The Wall Street Journal reported, only 20 days after Gemini 3.7 Flash arrived. The Journal also reported that Google employees preferred the new model, internally called Skimaki, to an unspecified Anthropic Opus model in head-to-head work inside Google’s Jetski coding tool. The claim would be consequential for coding-model buyers, but it has not yet been tested in public.

The reported release would extend an unusually short update cycle for Google’s Flash line. Gemini 3.7 Flash launched on August 13, and Google set introductory prices of $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. Those rates set a concrete reference point for developers evaluating whether a successor preserves Flash’s intended economics for high-volume work.

What the internal test can show

An engineer choosing one model over another while completing work can be a useful product signal. But the reported Jetski result does not establish that Gemini 3.8 produces more correct, secure, or maintainable code across repositories and development environments. The tool, evaluators, and test population were all within Google, and familiarity with Gemini’s interface, speed, instruction style, or integration could affect preference.

The public baseline is mixed

Google’s published figures for Gemini 3.7 Flash show why an internal preference should not be treated as a settled market ranking. On Google’s FrontierCode 1.1 evaluation, 3.7 scored 43.6%, compared with 42.7% for Claude Sonnet 5. Yet Google’s own model-card comparisons showed 3.7 trailing OpenAI’s GPT-5.6 Terra on DeepSWE and Terminal-bench software-engineering evaluations.

That contrast is central to the Gemini 3.8 story. A coding model can look stronger on one evaluation or feel better to engineers in a particular tool without taking a consistent lead across software tasks. For teams running agentic coding workflows, the eventual comparison will turn on task quality alongside latency and price, rather than an internal preference alone.

A faster release is not yet a product

If Gemini 3.8 Flash appears on the reported timetable, Google will have moved from one numbered Flash release to the next in 20 days. That is the immediate competitive signal: Google may be iterating its lower-cost coding line faster. Whether the iteration changes developers’ choices remains unresolved until the model is publicly available with comparable performance, safety, and pricing information.