DeepSeek has begun shipping DeepSeek-V4-Pro-0813 through its API and chat app, charging $0.435 per million input tokens and $0.87 per million output tokens for a model the lab says beats Anthropic’s Opus 4.8 on four internal benchmarks. The pricing lands the same week DeepSeek, the Hangzhou-based lab whose V3 release shipped at a fraction of frontier training costs, was reported to have trailed only Anthropic in monthly token volume for July. That ranking, drawn from a token-volume estimate the analyst known as tphuang posted on X and that Wccftech cited, would put DeepSeek’s inference volume ahead of OpenAI and Google for the month, a claim worth separating clearly from the benchmark marketing that came with the new model.
Start with what DeepSeek controls directly: the price. V4-Pro-0813 carries a 1.6 trillion total parameter count with 49 billion active per token and a 1 million token context window, according to a post from the coding tool Cline, which also claimed a 15.8 percent gain on Terminal Bench over DeepSeek’s April preview model. DeepSeek says the new model also outperforms Opus 4.8 on Cybergym, DeepSWE, and AutomationBench. Those results are preliminary figures circulating on WeChat, generated and selected by DeepSeek itself. Anthropic has not published a response, and DeepSeek has not disclosed Opus 4.8’s own pricing for a direct cost-per-performance comparison, so the benchmark claims stand as marketing until an independent lab reruns them.
The pricing context matters more than the benchmark claims. Days earlier, OpenAI cut GPT-5.6 Luna’s rates by as much as 80 percent, dropping input tokens from $1 to $0.20 per million and output tokens from $6 to $1.20 per million. Hours after that cut, DeepSeek shipped a refreshed Flash-class model, V4-Flash-0731, at just 284 billion parameters and priced at $0.14 per million input tokens and $0.28 per million output tokens, undercutting OpenAI’s discounted rate on both ends.
V4-Pro-0813 sits in a different spot on that grid. Its input price of $0.435 per million tokens runs more than twice OpenAI’s discounted $0.20, while its output price of $0.87 per million tokens runs about 28 percent below OpenAI’s $1.20. The arithmetic favors whichever workload dominates a given product. Agentic coding tools and long-form generation lean output-heavy, where V4-Pro comes out cheaper. Retrieval-heavy pipelines that stuff large context windows into every call lean input-heavy, where OpenAI’s cut rate wins.
For a US or EU buyer deciding where to route inference, price is one input among several. DeepSeek’s compute base is reported at roughly 20,000 Nvidia H100 GPUs, a fraction of what Anthropic, OpenAI, or Google run, and users have posted about inference slowdowns during peak hours that would track with that constraint. Routing production traffic through a Chinese lab’s API also raises data-residency and export-control questions that a domestic vendor does not, regardless of the per-token math. Neither factor shows up in a pricing table, but both belong in the same spreadsheet as the token costs before a contract gets signed.
The July token-volume figure deserves more weight than the benchmark wins, because it measures what customers actually did rather than what DeepSeek chose to test. If DeepSeek holds or extends that position through August, as the cited estimate speculates it might, the company’s pricing strategy will have converted into real market share faster than its benchmark claims can be independently checked. Teams evaluating inference vendors this quarter should model their own input-to-output token ratio against both DeepSeek’s and OpenAI’s new rate cards before assuming either one is cheaper by default.
Wccftech (Rohail Saleem) reported this story on August 12, 2026.