Qwen Releases Qwen3.8-Flash With 1M-Token Option and One-Ninth Training-Cost Claim
The new model gives developers a published context limit and usage price, while Qwen’s performance and efficiency comparisons remain its own stated results.
Loading page…
The new model gives developers a published context limit and usage price, while Qwen’s performance and efficiency comparisons remain its own stated results.
Listen to this story
Alibaba’s Qwen is pairing a paid API launch with an open-weight preview of its next architecture. Qwen3.8-Flash supports 262,144 tokens by default and up to 1 million, with API pricing of ¥1 per million input tokens and ¥3 per million output tokens. Alibaba positions it as a coding and office-task upgrade over Qwen3.7-Plus and claims roughly one-ninth its training cost, but provides no shared benchmark in this release.
Released August 26, Qwen3.8-Flash expands from a 262,144-token default context to 1 million tokens.
API pricing is ¥1 per million input tokens and ¥3 per million output tokens.
The one-ninth training-cost figure is Alibaba’s claim, not a result from a common benchmark.
Qwen has released a long-context multimodal model with a clear commercial offer: API access priced by token volume and a context window that can reach 1 million tokens. Alibaba also says Qwen3.8-Flash improves coding and office tasks while requiring about one-ninth the training cost of Qwen3.7-Plus—comparisons supplied by the company.
Alibaba’s Qwen released Qwen3.8-Flash on August 26 as a multimodal model, meaning it can work across more than one type of input. Its default context window is 262,144 tokens, expandable to 1 million. Qwen says the larger window is meant to accommodate large files, lengthy conversations and extensive research materials within an interaction.
The company positions the model as an upgrade from Qwen3.7-Plus for coding and office tasks. Its stated cost comparison is equally striking: Alibaba says Qwen3.8-Flash requires roughly one-ninth of the earlier model’s training cost. The announcement does not turn those assertions into a common benchmark, so they should be read as Alibaba’s product claims rather than a cross-model ranking.
Qwen’s API schedule charges 1 yuan per million input tokens and 3 yuan per million output tokens. Input tokens are the material sent to a model; output tokens are the generated response. Publishing separate rates makes the cost structure legible for applications that both submit long source material and generate substantial output.
The paired releases split Qwen’s message between a model available through a paid API and an open-weight technical preview. Qwen3.8-Flash offers declared operating limits and prices today. Flash-Next gives developers a route to examine the design Qwen says will underpin its next model family, without establishing a final Qwen4 specification.
Editorial analysis
This is a two-part release. Qwen3.8-Flash sets terms that developers can use immediately: a context window, an expansion ceiling and API prices. Flash-Next serves a different purpose. By releasing its weights, Qwen is inviting scrutiny of an architecture it says will inform Qwen4. The important next signal is whether developer assessment of Flash-Next clarifies the trade-offs behind Qwen’s efficiency pitch before that next model family arrives. Alibaba’s broader AI spending makes the architecture preview more than a routine companion release.
Loading discussion...
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.