Modelspublished

Qwen Releases Qwen3.8-Flash With 1M-Token Option and One-Ninth Training-Cost Claim

The new model gives developers a published context limit and usage price, while Qwen’s performance and efficiency comparisons remain its own stated results.

By 3 min read
Qwen Releases Qwen3.8-Flash With 1M-Token Option and One-Ninth Training-Cost Claim

Listen to this story

The audio brief

About 1:26
0:001:26
Read transcript
Qwen has launched Qwen3.8-Flash with a context window that can stretch to 1 million tokens, alongside an API price of just 1 yuan per million input tokens and 3 yuan per million output tokens. The model is multimodal, and its default context is 262,144 tokens before that expansion is enabled. In practical terms, developers can send in very large files, long conversations, or substantial research material in one interaction, then pay separately for what they provide and what the model generates. Alibaba positions Qwen3.8-Flash as an upgrade over Qwen3.7-Plus for coding and office work. It also says the new model requires about one-ninth of the earlier model’s training cost. That figure is Alibaba’s own comparison, though—not the outcome of a shared benchmark—so it describes the company’s claim rather than an independently established ranking. There’s a second release aimed at developers who want to look further ahead. Qwen3.8-Flash-Next comes with open-source weights and is presented as a prototype for the Qwen4 architecture. That creates a clear split: Flash is the commercially available model with published limits and pricing, while Flash-Next is a technical preview that can be inspected and evaluated. The key constraint is that this preview does not establish final Qwen4 specifications. For now, the concrete offer is the million-token option and its low, clearly stated API rates.

Story brief

3 key points

Alibaba’s Qwen is pairing a paid API launch with an open-weight preview of its next architecture. Qwen3.8-Flash supports 262,144 tokens by default and up to 1 million, with API pricing of ¥1 per million input tokens and ¥3 per million output tokens. Alibaba positions it as a coding and office-task upgrade over Qwen3.7-Plus and claims roughly one-ninth its training cost, but provides no shared benchmark in this...

  1. 01

    Released August 26, Qwen3.8-Flash expands from a 262,144-token default context to 1 million tokens.

  2. 02

    API pricing is ¥1 per million input tokens and ¥3 per million output tokens.

  3. 03

    The one-ninth training-cost figure is Alibaba’s claim, not a result from a common benchmark.

Qwen has released a long-context multimodal model with a clear commercial offer: API access priced by token volume and a context window that can reach 1 million tokens. Alibaba also says Qwen3.8-Flash improves coding and office tasks while requiring about one-ninth the training cost of Qwen3.7-Plus—comparisons supplied by the company.

Alibaba’s Qwen released Qwen3.8-Flash on August 26 as a multimodal model, meaning it can work across more than one type of input. Its default context window is 262,144 tokens, expandable to 1 million. Qwen says the larger window is meant to accommodate large files, lengthy conversations and extensive research materials within an interaction.

The company positions the model as an upgrade from Qwen3.7-Plus for coding and office tasks. Its stated cost comparison is equally striking: Alibaba says Qwen3.8-Flash requires roughly one-ninth of the earlier model’s training cost. The announcement does not turn those assertions into a common benchmark, so they should be read as Alibaba’s product claims rather than a cross-model ranking.

Qwen’s API schedule charges 1 yuan per million input tokens and 3 yuan per million output tokens. Input tokens are the material sent to a model; output tokens are the generated response. Publishing separate rates makes the cost structure legible for applications that both submit long source material and generate substantial output.

A second release points ahead

  • Qwen released open-source weights for Qwen3.8-Flash-Next.
  • Qwen says Flash-Next is a prototype for the next-generation Qwen4 family and is intended to let developers evaluate that architecture.

The paired releases split Qwen’s message between a model available through a paid API and an open-weight technical preview. Qwen3.8-Flash offers declared operating limits and prices today. Flash-Next gives developers a route to examine the design Qwen says will underpin its next model family, without establishing a final Qwen4 specification.

Editorial analysis

Our Read

Our read: This is a two-part release. Qwen3.8-Flash sets terms that developers can use immediately: a context window, an expansion ceiling and API prices. Flash-Next serves a different purpose. By releasing its weights, Qwen is inviting scrutiny of an architecture it says will inform Qwen4. The important next signal is whether developer assessment of Flash-Next clarifies the trade-offs behind Qwen’s efficiency pitch before that next model family arrives. Alibaba’s broader AI spending makes the architecture preview more than a routine companion release.

Sources

  1. yahoo.comAlibaba's Qwen launches Qwen3.8-Flash AI model with lower training costs
  2. m.economictimes.comAlibaba's Qwen launches Qwen3.8-Flash AI model with lower training costs - The Economic Times