DeepSeek V4 Goes GA: 1.6T Parameters, 1M Context, Input Pricing as Low as $0.14/M Tokens
DeepSeek V4 hit general availability on July 20, and legacy API endpoints were permanently shut down on July 24 — the Chinese AI lab has officially entered mass-production mode.
Three Key Takeaways
Flagship model pushes parameter ceilings. DeepSeek-v4-pro packs 1.6 trillion total parameters (49B active) with a 1-million-token context window and up to 384K output tokens. The GA build adds stronger math reasoning, code generation, and agentic capabilities, putting it in direct competition with GPT-5.6 and Claude's flagship offerings.
The price war escalates — but with a twist. Input pricing drops to $0.14/M tokens for v4-flash and roughly $0.70/M for v4-pro. However, Beijing-time hours of 9–12 and 14–18 are designated as "peak," with prices doubling. This marks the first time a major LLM provider has adopted a utility-style "time-of-use" pricing model — essentially using price signals to manage GPU load.
Legacy APIs were cut off on July 24. The old model names deepseek-chat and deepseek-reasoner stopped working at 15:59 UTC on July 24. All users had to migrate to V4's new endpoints. DeepSeek gave about three months of migration runway, but for developers still on the old API, it was a forced bloodletting.
WangDou's Take
The most interesting part of DeepSeek's move isn't the parameter count — it's the peak-hour pricing. A floor price of $0.14/M tokens is genuinely aggressive — OpenAI charges twenty to thirty times more for comparable models at standard rates. But the peak-time doubling reveals a hard truth: even DeepSeek doesn't have enough GPUs. Rather than letting everyone pile onto the same time slots and melt the cluster, they're using price to spread demand. That's actually smarter than a mindless price war — you can use it cheap, but you have to go off-peak. Meanwhile, the 1.6T-parameter, 49B-active MoE architecture means inference costs are structurally lower than dense models — that's the real foundation for the low prices. As for cutting off legacy APIs with zero sympathy? Classic DeepSeek: keep up or get left behind.
Source: Tech Insider, SitePoint, ExplainX
