DeepSeek started retiring its flagship model by rerouting every request to a cheaper one
From 04:00 UTC on September 14, every request to DeepSeek's V4-Pro model routes to V4.1-Flash. Requests are billed at V4.1-Flash prices. The company says the switch holds until V4.1-Pro launches. It says the older V4-Flash and V4-Flash-Vision-Exp models are already retired.
Verified 7:20 AM PDT · 3 original sources
V4.1-Flash is a 552-billion-parameter mixture-of-experts model. It runs on a new causal encoder-decoder architecture. The model has 8 billion active parameters for input and 16 billion for output. DeepSeek says its cache needs a quarter of the high-bandwidth memory of the previous generation. It also needs an eighth of the solid-state storage. DeepSeek says cache charges are a large share of agent costs.
The company says tests by several parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed and total runtime. New pricing took effect on September 10. Reuters reported the launch alongside the company's preparation for an initial public offering on Shanghai's STAR Market.
Every performance figure is the vendor's own until an outside party reproduces it on the public API. Caching savings depend on workloads, and a cheaper flagship cuts the revenue that pays for the next training run. The efficiency push is also a way around limits on hardware. Export controls have limited access to Nvidia accelerators and high-bandwidth memory, which raises the price of Chinese substitutes.
Watch for independent runs of the agent benchmarks on the public API and for measured cache costs in production. The next two events are the V4.1-Pro launch and the STAR Market filing. Both would put the company's spending and margins on the record.
Audit the story
Original sources
Company claims remain company claims. Follow the reporting and judge the evidence directly.
Continue the edition