DeepSeek Launches V4.1-Flash and Plans to Replace V4-Pro API Traffic

The company says its smallest model in a new architecture family improves speed, cost and total runtime—but its own testing is driving the planned V4-Pro handoff.

By 3 min read
DeepSeek Launches V4.1-Flash and Plans to Replace V4-Pro API Traffic
DeepSeek Launches V4.1-Flash and Plans to Replace V4-Pro API Traffic

Listen to this story

The audio brief

About 1:25
0:001:25
Read transcript
DeepSeek is planning to turn its new V4.1-Flash model into the automatic replacement for customers who chose the higher-tier V4-Pro API. After noon Beijing time on September fourteenth, requests sent to the deepseek-v4-pro endpoint are expected to route to V4.1-Flash instead, and customers will pay the Flash rate until V4.1-Pro arrives. That makes this more than a routine model launch. V4.1-Flash is the smallest model in DeepSeek’s new architecture family, but it adds native visual input and is designed to improve speed, throughput, and the ability to scale toward larger models. DeepSeek says its own testing put Flash ahead of V4-Pro on performance, cost, speed, and total runtime. Those are company-reported results, though, so they may not hold across every customer workload. The model is available through the API as deepseek-flash. For compatibility, the older deepseek-v4-flash and deepseek-v4-flash-vision-exp names will temporarily point to it, while the previous V4-Flash and V4-Flash-Vision-Exp models have been retired. DeepSeek has also cut API prices, with the detailed rates listed separately. Both V4-Flash and V4-Pro were introduced with million-token context windows. The immediate question is whether customers who deliberately selected Pro find this automatic Flash handoff suitable before V4.1-Pro is released.

Story brief

3 key points

DeepSeek is making V4.1-Flash the likely default for customers using its higher-tier V4-Pro API: after 12:00 Beijing time on September 14, deepseek-v4-pro requests are planned to route to Flash and use Flash pricing until V4.1-Pro arrives. The new model adds native visual input and is positioned as the smallest V4.1 architecture model. DeepSeek claims it beats V4-Pro on tested performance, cost, speed, and total...

  1. 01

    New calls can use deepseek-flash; legacy deepseek-v4-flash and deepseek-v4-flash-vision-exp names will temporarily preserve compatibility.

  2. 02

    DeepSeek retired the prior V4-Flash and V4-Flash-Vision-Exp models and reduced API prices; detailed rates are on a separate pricing page.

  3. 03

    V4-Flash and V4-Pro were introduced with 1-million-token context windows, preserving that family-level positioning in the comparison.

DeepSeek has released V4.1-Flash, a new AI model with native visual understanding, and says it will soon direct requests from its V4-Pro API endpoint to the new model. The move puts a smaller member of DeepSeek’s new architecture family in line to become the default for customers who had selected the company’s higher-tier option.

The release is available through the DeepSeek API as deepseek-flash. DeepSeek calls V4.1-Flash the smallest model in its new architecture family and says the design is intended to raise the model’s capability ceiling while improving inference speed, throughput and the ability to scale to larger models. Native multimodal support means the model can accept visual as well as text input.

One new model, two different promises

For prospective users, DeepSeek’s first promise is a conventional product claim: a faster, more capable and higher-throughput Flash model. The second is more consequential. The company says extensive testing found V4.1-Flash ahead of V4-Pro on performance, cost, speed and total time. That is DeepSeek’s assessment of its own tested workloads, rather than an independent ranking of the two systems.

A migration, not just another endpoint

The switchover changes the practical meaning of the launch for existing API customers. Rather than asking every V4-Pro user to choose a successor, DeepSeek plans to make V4.1-Flash the destination automatically until it releases V4.1-Pro. The planned switch trades continued access to a named higher-tier endpoint for DeepSeek’s claim that its Flash replacement is the stronger deal.

What API users need to know

  • New calls can use deepseek-flash to reach V4.1-Flash.
  • The older deepseek-v4-flash and deepseek-v4-flash-vision-exp names temporarily point to V4.1-Flash for compatibility.
  • DeepSeek has retired the prior V4-Flash and V4-Flash-Vision-Exp models.

Flash moves closer to the flagship

That positioning sharpens a contrast already present in the V4 family. When DeepSeek introduced V4 in April, it described V4-Flash as the faster, more economical alternative and V4-Pro as the model aimed at the highest performance. Both models were presented with a 1-million-token context window. V4.1-Flash now carries the Flash label, but DeepSeek says it can surpass V4-Pro across the measures that matter for its planned replacement.

DeepSeek has also reduced API prices with the release, though its public changelog directs users to a separate pricing page for the details. Its headline comparisons of V4.1-Flash with V4-Pro should therefore be read as company-supplied performance and cost claims. The announcement establishes DeepSeek’s product direction and migration plan; it does not by itself establish how the two models will compare on every customer workload.

The launch arrives as DeepSeek prepares for a potential initial public offering on Shanghai’s technology-focused STAR Market, according to Reuters. For now, the immediate test is less about that prospective listing than the forced comparison embedded in the API: whether customers who chose V4-Pro find the new Flash default suitable once the routing change begins.

Editorial analysis

Our Read

DeepSeek’s consequential move is the planned endpoint handoff, not merely the new model name. A company can offer a lower-cost fast model alongside a flagship without forcing a choice. Routing V4-Pro traffic to V4.1-Flash makes DeepSeek’s internal conclusion about quality, speed and cost an operating decision for API customers. The next revealing event is the September 14 cutover: whether DeepSeek keeps that schedule and how developers respond will show whether a single cheaper default can satisfy workloads that previously used a higher-tier model. The company’s recent V4 releases have already made fast, economical models a central part of the family’s pitch.

Sources

  1. api-docs.deepseek.comChange Log | DeepSeek API Docs
  2. deepseek.comDeepSeek-V4 Preview: Entering the Era of Affordable Million-Token Context
  3. thehindubusinessline.comDeepSeek launches V4.1-Flash AI model ahead of Shanghai IPO

Loading discussion...