On the evening of July 31, 2026, DeepSeek released the official version of DeepSeek-V4-Flash (checkpoint 0731) in the changelog of its own API documentation. No launch event, no teaser posters, not even a formal announcement article — just a few paragraphs of notes in an update log.
A year and a half earlier, on the Monday the same company released its R1 model, Nvidia fell 17% in a single day, erasing 589 billion dollars of market value — the largest one-day market-cap loss in US stock market history. Global financial media called that day the "DeepSeek moment." This time, on Friday, July 31, 2026, Nvidia closed up 2.93% at $200.75, and the Nasdaq closed up 1.0% — not a single major financial outlet connected the day's trading to this release.
It looks like "no impact." But zoom out, and this calm is itself the most important change of the past eighteen months: breakthroughs from Chinese large models have gone from black-swan events to the market's default expectation. Shock only happens once; the norm is what defines the landscape.
This report reviews the release across three coordinate systems: the engineering meaning of the model itself, the real reaction of global capital markets, and the deeper shifts in the US-China AI competition. Finally, it answers the question the Tianxia Gongchang Industry Research Institute cares about most — what does the continued slide in inference costs mean for Chinese manufacturing?
1. The Release Itself: An Engineering Philosophy of the Underdog Overtaking
Start with what this release actually delivered.
V4-Flash-0731 is a mixture-of-experts (MoE) model with 284 billion total parameters, only 13 billion of which are activated per inference, a context length of 1 million tokens, and fully open weights under the MIT license — commercial use, fine-tuning, and redistribution are all unrestricted. The official changelog notes that its architecture and size are identical to the V4-Flash preview released in April; it was "only re-post-trained."
The report card is what matters. By the official account, this "small" model with 13 billion active parameters overtook its own sibling — the V4-Pro preview with 1.6 trillion total and 49 billion active parameters — across all nine agentic benchmarks: Terminal Bench 2.1 jumped from the preview's 61.8 to 82.7, with repository understanding, cybersecurity, and full-stack development benchmarks all punching above its weight class. And when the V4-Pro preview launched in April, the official wording was that it "approaches or even matches" the top international flagship models. Chinese tech media gave this release an accurate name: "the underdog overtaking from below" (Zhidx's launch coverage).
Third-party evaluation corroborates the official numbers. The independent benchmark firm Artificial Analysis gave V4-Flash-0731 an Intelligence Index score of 50 — third in intelligence among the 101 comparable models it tracks, while ranking only 22nd on price; the median Intelligence Index of comparable models in its price band is just 17 (Artificial Analysis model page). Two comparisons make this concrete. Longitudinally: April's preview scored 40 on the same index, so a single post-training iteration added 10 points. Horizontally: an absolute score of 50 sits in the same band as GPT-5.6 Luna's 51 and Gemini 3.6 Flash's 50 — for the first time, a Chinese open-source model has taken a seat in the top tier of lightweight models.
Pricing did not change: 1 yuan per million input tokens and 2 yuan per million output tokens, with cache hits as low as 0.2 yuan — roughly $0.14 / $0.28, about one-tenth the price of the company's own V4-Pro.
There is a detail here that every engineering organization should remember: this performance leap did not change a single line of architecture code. All of the gains came from post-training — iterations on data, methods, and training strategy. It means that beyond the pre-training scale race there exists another steep curve of capability improvement, and this curve consumes far less compute than stacking parameters. For a Chinese company whose training compute faces external constraints, the strategic meaning of this route speaks for itself.
Another detail is just as intriguing: V4-Flash-0731 is the first DeepSeek model to natively support OpenAI's Responses API format, with dedicated adaptation for the Codex coding tool (Sina Tech's report). In other words, DeepSeek plugged itself directly into its competitor's developer ecosystem — using the rival's interface standard to serve the rival's users, driving switching costs down to nearly zero.
2. Why the Market Barely Moved
Start with the closing data for Friday, July 31 (US individual stock data from the historical quote pages of stockanalysis.com).
All three major US indices closed higher: Nasdaq +1.0%, S&P 500 +0.7%, Dow +0.53%, with the Dow logging its fourth straight monthly gain. But the day's protagonists were written plainly on the tape — Amazon surged 15.32% after beating earnings expectations, Alphabet rose 6.73%, Microsoft 3.02%, Meta 3.28%. This was a trading day dominated by Big Tech earnings.
The semiconductor sector's path is even more telling. The Philadelphia Semiconductor Index rose more than 5% intraday, gave nearly all of it back by the close, and finished up just 0.07%; Nvidia closed up 2.93% at $200.75, AMD fell 1.90%, and Micron — after soaring 18.36% the previous day — pulled back 5.90% (Sina Finance's evening wrap recorded that intraday plunge). One thing deserves special emphasis: the real semiconductor surge came a day earlier, on July 30, when Microsoft's earnings ignited an 8.19% single-day jump in the SOX; July 31 was merely the digestion day after that spike. In other words, the week's semiconductor action was driven by US giants' capex guidance, not by DeepSeek — a search across major English-language financial media turns up not one article linking July 31's trading to V4 Flash.
Against January 2025, the contrast is dramatic. For an equally major DeepSeek release, R1 knocked Nvidia down 17% in a day and erased $589 billion; on the day the official V4 Flash shipped, the market did not give it so much as a proper red candle. There are three layers of reasons.
First, the "low-cost shock" has been fully priced in. CNBC ran a piece at the start of this year analyzing exactly this: January 2025 triggered a global repricing because R1 changed, for the first time, the market's understanding of the frontier-model cost curve and of Chinese competitiveness; every DeepSeek release since has lacked that "cognition-shattering" property (CNBC's start-of-year retrospective). When the V4 preview launched this April, the market already shrugged — consulting analysts put it as "the wow factor happened last year; it is long since in the price."
Second, the capex narrative has inverted. The panic logic of the R1 moment was "cheaper models mean the giants won't need as many chips"; eighteen months later, the market has accepted the opposite framework — efficiency gains drive usage explosions, economics' Jevons paradox. In 2026, Meta, Microsoft, Amazon, and Alphabet are expected to spend roughly $475 billion in combined capex, not a cent of it revised down because of Chinese models' cost reductions. On release day, Bloomberg carried another story: Moonshot AI locked in roughly 20,000 Nvidia chips through an Alibaba Cloud contract — Chinese model companies are racing to acquire compute, not economize on it.
Third, the release itself was restrained. The official version covers only the Flash tier; the flagship V4-Pro official release is still pending. Guancha.cn's headline captured it precisely: "The official V4 is here — but only half of it" (Guancha's report). The market naturally saved its biggest suspense for the next act.
China's A-shares deserve a separate note. On July 31, the Shanghai Composite closed up 0.72% at 3,832.26, the Shenzhen Component rose 2.21% to 13,578.93, and the ChiNext Index jumped 3.06% to 3,343.96, with combined turnover above 2.5 trillion yuan, more than a hundred stocks limit-up and zero limit-down; the compute hardware chain rallied across the board — Cambricon up 6.10%, Foxconn Industrial Internet up 5.39%, SMIC's A-shares up 0.81%. But the attribution was mostly not DeepSeek — the same day, the National Development and Reform Commission stated that the national integrated computing-power network would add 4 trillion yuan of new direct investment during the 15th Five-Year Plan, which, layered on an end-of-month recovery after July's deep correction, was the main driver (Sina Finance's closing report). Hong Kong was comparatively quiet, with the Hang Seng roughly flat on the day and up more than 13% for July. The V4 Flash news trickled out through an API changelog in the evening; A-shares' independent pricing would wait for the following week. Weekend brokerage strategy notes confirmed the "normalization" reading: China Galaxy Securities' strategy team sees August entering a phase of "dual verification by policy and earnings," with the AI trade shifting from broad rallies to earnings-based selection, while Zhongtai Securities maintains that "technology remains the main line" (Sina Finance's August 1 strategy roundup) — not one institution treated the model release itself as a trading signal. AI has gone from an "event" to a "sector cycle."
3. The Price War Flips: This Time OpenAI Cut First
If the market reaction was "nothing happened," the industrial competition was the opposite — in the 48 hours around this release, something happened whose sequence runs against intuition.
On July 30, the day before the official V4 Flash shipped, OpenAI announced an 80% price cut on GPT-5.6 Luna — released just 21 days earlier — dropping input pricing from $1 to $0.20 per million tokens (VentureBeat's analysis of the cuts). A flagship-derived model discounted to one-fifth of its price three weeks after launch has no precedent in OpenAI's history.
Where the pressure comes from, the data states bluntly: CNBC's July survey showed Chinese models already account for 46% of US enterprise token usage on the OpenRouter platform; another tally puts Chinese open-source models' share of global traffic at 66.5%, outrunning US models for 13 consecutive weeks (36kr's overseas edition on the statistics). In the market where developers vote with their feet, offense and defense have long since traded places.
More interesting still is DeepSeek's response: no price cut. The official V4 Flash held its price and answered with capability upgrades and ecosystem compatibility — performance overtaking its own flagship preview, interfaces plugged straight into the rival's ecosystem. Third-party benchmark costing shows that even after Luna's 80% cut, completing a single agentic task on V4 Flash 0731 still costs roughly 60% less than on the competitor (IT Home's benchmark costing). When your rival is forced to cut prices and you can afford not to, it is obvious who holds the initiative in a price war.
Industry commentary sums up this phase as "intelligence is becoming as cheap as water and electricity." Capability itself is no longer the moat; cost structure is — and cost structure happens to be the battlefield where China's engineering system excels. Chinese manufacturing has run this playbook repeatedly in photovoltaics, lithium batteries, and new-energy vehicles; now, for the first time, it is replaying in frontier AI.
4. US-China AI Competition: Algorithmic Efficiency Is Rewriting the Rules
Shift the lens from markets to geo-technological competition, and the V4 Flash release is not an isolated event but a footnote added to each of three evolving storylines.
Storyline one: algorithmic efficiency is partially offsetting compute constraints. The underlying logic of export controls is to use a generational gap in advanced compute to lock in a generational gap in model capability. Yet with one pure post-training iteration, V4 Flash pushed a 13-billion-active-parameter model past the 49-billion-active flagship preview — intelligence per unit of compute is still climbing steeply. Third-party measurements show the V4 series needs only 27% of the per-token inference compute of the previous-generation V3.2 in million-token long-context scenarios. When efficiency improves faster than compute restrictions tighten, the marginal effectiveness of the blockade decays.
Storyline two: the loosening of the software ecosystem cuts deeper than the chips themselves. The V4 technical report — unusually — carried out native engineering validation on both Nvidia GPUs and Huawei Ascend NPUs, and Huawei disclosed that Ascend chips took part in portions of V4-Flash's training and received full adaptation support (TrendForce's analysis of the Ascend adaptation). Nvidia founder Jensen Huang, in a podcast interview, called the situation in which leading models no longer default to CUDA as their optimization starting point "catastrophic" — a word that carries weight coming from the biggest beneficiary of the CUDA ecosystem. The accompanying industrial fact: Ascend's total 2026 production capacity is planned at roughly 1.6 million units, giving domestic compute, for the first time, the supply capability to absorb frontier models at scale. The absorption speed proves the point: on August 1, the day after release, xFusion's AI Lab announced V4-Flash-0731 support; on August 2, the National Supercomputing Internet platform put its API online for one-click access (Sina Tech's follow-up); and Huawei announced that its entire in-market Ascend lineup supports both V4-Flash and V4-Pro. The interval from model release to full availability on domestic compute platforms has compressed to a matter of days.
Storyline three: America's internal split over "blocking open source" has gone public. In the week before the release, the US AI industry's debate over whether to restrict Chinese open-source models reached a crescendo: Meta founder Mark Zuckerberg said publicly that a ban "would not be an effective solution"; Anthropic CEO Dario Amodei, while voicing safety concerns about open-weight models, was equally clear in opposing a ban, calling one "protectionist and ineffective." The Register's commentary headline was blunter still — "The truth nobody wants to admit: open models are competitive now" (The Register's industry commentary). The Atlantic Council's policy brief reminded Western readers: "The best AI you can own is Chinese." When MIT-licensed open weights, near-zero switching costs, and an order-of-magnitude price advantage stack together, there is indeed little that administrative measures can change.
It should be noted that as of this writing, Washington has made no official statement on the 0731 release specifically — it was, after all, just an API iteration. What could genuinely trigger the next round of policy reaction is the upcoming official V4-Pro and whether it discloses its training hardware.
5. What This Means for Chinese Manufacturing
When the Tianxia Gongchang Industry Research Institute follows a release like this, the ultimate concern is not the capital markets but a plain industrial question: when agent-grade AI capability costs one or two yuan per million tokens, what can manufacturing do with it?
For the past two years, what held back AI adoption in manufacturing B2B scenarios was never model capability — it was unit economics. Using large models to analyze supplier qualifications company by company, clean enterprise data record by record, or parse tender documents clause by clause has long been technically feasible; at flagship-model prices, one look at the arithmetic ended the conversation. What the V4 Flash generation changes is precisely that arithmetic: with per-task agent costs entering the "cents" range, AI due diligence, AI factory sourcing, and AI supply-chain analysis — once affordable only to head enterprises — for the first time have a cost structure that can reach small and medium factories and trading companies.
This is also the work we do. The Tianxia Gongchang platform indexes 4.8 million Chinese manufacturing factories screened for authenticity, and its AI factory-search service lets buyers and sellers complete supplier screening that used to take days with a single natural-language sentence. The inference cost supporting this kind of service has fallen by an order of magnitude over the past twelve months — and this release tells us the cost curve is nowhere near its end.
For Chinese manufacturing, this is a rare structural opportunity: the world's cheapest frontier AI capability, the world's most complete manufacturing system, and a domestic compute foundation taking shape — appearing in the same coordinate system for the first time. The last time a comparable combination of factors appeared, China spent a decade driving the levelized cost of photovoltaic power down to a level that made the whole world redo its energy planning.
Conclusion: Quiet Is the Sound of Maturity
Looking back on this release, what most deserves recording is not any benchmark score but the unprecedented quiet — the publisher was quiet, the market was quiet, and even the reaction across the Pacific was measured.
The storm of January 2025 came from the world realizing, for the first time, that a Chinese team could build a frontier model at one-tenth the cost; the calm of July 2026 shows that this no longer needs to be "realized" at all. From outlier to baseline — that is the finest coming-of-age a technology route can have.
The dates to watch next: the official V4-Pro release — expected in early August — and its training-hardware disclosure; A-share compute and AI-application sectors' independent pricing of this model cycle; and the chain reaction across the global API price system after OpenAI's 80% cut. The Tianxia Gongchang Industry Research Institute will keep tracking.
Cover image: the Tianhe-2 supercomputer machine hall, photographed by O01326, via Wikimedia Commons, licensed CC BY-SA 4.0.