DeepSeek’s newest budget model, V4 Flash “0731,” scored 50 points on the Artificial Analysis Intelligence Index, a single point behind OpenAI’s GPT-5.6 Luna, while charging developers roughly 60 percent less to run a comparable task, according to The Decoder. That price gap survives OpenAI’s own decision to cut Luna’s rates by 80 percent earlier this month. Two labs are no longer competing mainly on what their models can do. They are competing on what running them costs.
The pattern is now familiar. OpenAI moved first, slashing prices on its cheapest GPT-5.6 tier to hold onto developers who route high-volume workloads to whichever model is good enough and cheapest. DeepSeek, the Hangzhou-based lab known for shipping capable models at a fraction of typical training and inference cost, answered within the same month by closing most of the remaining quality gap and beating the new, lower OpenAI price anyway.
Part of the cost advantage comes from infrastructure, not just the model itself. DeepSeek’s published API pricing offers a 98 percent discount on cached tokens, above the roughly 90 percent that has become standard among competing providers. V4 Flash 0731 also uses 12 percent fewer tokens per task than the version it replaces, so the lower per-token rate and the leaner token usage compound.
The performance gains behind the new score come from Artificial Analysis, an independent evaluator, rather than from DeepSeek’s own marketing copy, which is a meaningfully stronger form of evidence than a lab grading its own release. The index shows V4 Flash 0731 improving across every category it tracks compared with the April version, with the sharpest gains on agentic tasks. On GDPval, an Artificial Analysis benchmark that scores models against realistic office work, the model’s rating rose from 1,189 to 1,559 Elo points, and the index also recorded a lower hallucination rate than the prior release.
Still, an index score is not a production deployment. The Decoder’s report cites Artificial Analysis’s benchmark results but does not include independent testing of latency, uptime, or how the model performs on live agentic workloads at scale, only the aggregate score. Readers evaluating a switch should treat the 50-point figure as a strong signal on general capability, not proof that V4 Flash 0731 will hold up inside a specific production pipeline.
DeepSeek left the underlying architecture untouched: 284 billion total parameters with 13 billion active per token, and a one million token context window, per the model’s Hugging Face listing. The weights ship under an MIT license, DeepSeek’s own choice of terms, which lets any company deploy, fine-tune, or resell the model without paying DeepSeek a licensing fee.
The sequence matters more than either release alone. OpenAI cut its cheapest frontier-adjacent model by 80 percent, and within weeks a rival undercut that new price by another 60 percent while nearly closing the capability gap. That is two floor drops in the same pricing tier inside a single month, faster than either company’s own roadmap likely assumed when it set this year’s rates. Anyone building a product margin on top of GPT-5.6 Luna or a comparable DeepSeek tier should model per-token costs as a moving target for the rest of 2026, not a fixed input, and revisit vendor contracts that lock in current rates for more than a quarter.
The Decoder’s Thomas Joos reported on DeepSeek’s V4 Flash “0731” release and its Artificial Analysis Intelligence Index score on July 31, 2026.