DeepSeek has reversed its plan to retire the V4-Pro model on September 14, 2026, and says it will keep serving V4-Pro through its API with billing unchanged, citing user demand. The U-turn comes four days after the launch of DeepSeek-V4.1-Flash, a 552-billion-parameter mixture-of-experts model that DeepSeek prices at $0.15 per million input tokens and $0.60 per million output tokens during off-peak hours, less than a third of V4-Pro’s rate, and which the company claims outperforms its own flagship.
A four-day reversal
The original release notice on DeepSeek’s API docs, dated September 10, said V4-Pro was being phased out and that from 04:00 UTC on September 14 every deepseek-v4-pro request would be routed to V4.1-Flash and billed at Flash rates until a V4.1-Pro model launched. The Models and Pricing page now carries a footnote saying the company has decided to continue providing V4-Pro after September 14 in response to user demand, with the billing method unchanged, and that it will give further notice of any changes.
The two older Flash-generation models, V4-Flash and V4-Flash-Vision-Exp, are retired as planned. Their legacy model names still work, but requests are served by V4.1-Flash and billed at the Flash price. The new model name for the API is simply deepseek-flash.
Numbers behind the price cut
DeepSeek’s pricing page lists V4.1-Flash at $0.15 per million input tokens on a cache miss and $0.60 per million output tokens off-peak, doubling to $0.30 and $1.20 at peak. Cached input costs $0.003 per million off-peak and $0.006 at peak. V4-Pro, by comparison, is $0.66 input and $1.98 output off-peak, and $1.32 and $3.96 at peak. Both models support a 1 million token context window and up to 384K output tokens, but only V4.1-Flash supports vision. Flash also gets a concurrency limit of 2,500 versus 500 for V4-Pro.
The company says the price is possible because of a new Causal Encoder-Decoder architecture that activates only 8 billion parameters when reading input and 16 billion when generating output, out of 552 billion total. The KV cache, which is the memory a model uses to hold conversation context, needs a quarter of the high-bandwidth memory and an eighth of the SSD storage of the previous generation, and DeepSeek says cache-hit charges are often the largest share of an AI agent’s running cost. VentureBeat reported that on DeepSeek’s own benchmark runs, V4.1-Flash scored 74.2 on DeepSWE v1.1 against 74.0 for Claude Opus 5 and 73.0 for GPT-5.6 Sol, and that the weights are released on Hugging Face under an MIT licence. Those benchmark figures are DeepSeek’s, not independent, and should be read that way.
Why did users want V4-Pro back?
DeepSeek has not said. The most likely explanation is the ordinary one: production systems tuned to one model’s behaviour do not appreciate being silently switched to another, even a cheaper one that scores higher on benchmarks. Prompt formats, tool-calling habits and output style all shift between models, and a developer running a client’s customer-service bot would rather migrate on their own schedule. The retirement notice gave four days. Keeping V4-Pro alive, at a price several times higher than Flash, costs DeepSeek little and removes the pressure.
The peak-hour clock, in Pakistan and Gulf time
DeepSeek’s peak pricing applies from 01:00 to 04:00 and from 06:00 to 10:00 UTC, Monday to Friday. Everything else is off-peak at half price, including all weekend hours. In Pakistan Standard Time that makes peak hours 6 AM to 9 AM and 11 AM to 3 PM on weekdays. In the UAE and Oman, it is 5 AM to 8 AM and 10 AM to 2 PM. In Saudi Arabia, Qatar, Bahrain and Kuwait, it is 4 AM to 7 AM and 9 AM to 1 PM.
For a freelancer or small agency in Lahore or Dubai, that means the entire evening and night, when many Pakistani developers actually do client work for Western time zones, is already off-peak. Batch jobs such as bulk content generation, document processing or embedding runs can be scheduled after 3 PM PKT or on weekends and pay $0.15 rather than $0.30 per million input tokens. On a job that reads 100 million tokens, that is the difference between $15 and $30, before any cache savings.
Cost at this level stops being the thing that separates one provider from another. As beingguru noted last week, Upwork clients now search for AI roles, not AI tasks, and the skill being bought is judgement about which model to use where, how to keep it reliable, and how to handle data. On that last point, DeepSeek is a Chinese company, and clients in the EU, UK and the Gulf increasingly write data-residency terms into contracts. The open weights are the way around that: the MIT licence allows commercial self-hosting or use through a third-party host in whatever region the client requires.
What V4.1-Pro has to beat
DeepSeek’s original notice said the V4-Pro routing would last only until V4.1-Pro launched, so a larger model on the new architecture is coming. It now has to justify a premium over a Flash model that the company itself says is faster, cheaper and stronger than the current flagship. The V4-Pro reprieve buys developers time to migrate, but it also tells the market something: even at $0.15 per million tokens, the model that people have built on is the one they are reluctant to give up.





