Why this is noise: A permanent price cut on an API is a commercial pricing decision. The real story, China building a viable domestic AI stack outside U.S. semiconductor control — is already well established. This confirms the trend but does not change it.
Key Takeaways
- DeepSeek has made its 75% V4-Pro price cut permanent, dropping API costs to between 0.025 and 6 yuan per million tokens from 0.1 to 24 yuan previously
- ByteDance, Tencent, and Alibaba are all scrambling to secure Huawei Ascend 950 chip orders following V4’s release
- Huawei’s Ascend 950PR is the only domestic Chinese chip that supports the compressed numerical format required to run V4 efficiently
- Supply constraints remain, with Huawei targeting 750,000 units of the 950PR this year but production limited by U.S. export controls on chipmaking equipment
DeepSeek has permanently locked in a 75% price reduction on its flagship V4-Pro AI model, making costs that were introduced as a temporary promotional rate the new baseline for developers and enterprise users accessing the model via API.
V4-Pro now costs between 0.025 and 6 yuan per million tokens, down from 0.1 to 24 yuan, depending on usage type. DeepSeek did not disclose whether the permanent cut was enabled by improved supply of Huawei’s Ascend 950 chips, which the company uses to run V4 at scale.
When V4 launched last month, DeepSeek said the Pro version cost up to 12 times more than its less powerful Flash variant due to “constraints in high-end compute capacity.” It also said pricing was expected to fall sharply once Huawei Ascend 950 supernodes shipped in large quantities in the second half of 2026. The permanent cut arrives ahead of that timeline.
“Constraints would persist until production ramps up, reflecting the tight supply of high-end homemade AI chips.” — DeepSeek, via Reuters
The pricing move follows a scramble among China’s largest tech companies to secure Huawei chip supply triggered by V4’s release. ByteDance, Tencent, and Alibaba have all reached out to Huawei about new Ascend 950 orders, according to Reuters sources. Cloud platforms moved fast: Alibaba Cloud made V4 available on launch day, and Tencent Cloud launched preview services the same day across domestic nodes and its Singapore gateway.
The Ascend 950PR is currently the only domestic Chinese chip that supports the compressed numerical processing format V4 requires, giving Huawei a near-monopoly position inside China’s AI infrastructure stack at exactly the moment demand is surging.
V4-Pro at $0.0035–$0.83 per million tokens makes it one of the cheapest frontier-capable models available anywhere. For context, comparable Western models can run $3–$15 per million tokens at the input end, making V4’s permanent pricing a significant competitive signal for any enterprise evaluating API costs at scale.
Supply, however, remains the binding constraint. Huawei planned to ship around 750,000 units of the 950PR this year, with full-scale shipments beginning in the second half of 2026. U.S. export controls on advanced chipmaking equipment continue to cap how quickly Huawei can scale production, meaning demand is likely to outpace supply well into next year.
DeepSeek’s V4 models, V4-Pro with 1.6 trillion parameters and V4-Flash with 284 billion are both available as open-source releases under the MIT licence, allowing companies to freely use, modify, and commercialise them.
That combination of open weights, aggressive API pricing, and domestic chip optimisation represents the clearest signal yet that China’s AI stack is no longer dependent on U.S. infrastructure to reach competitive performance.
The piece being tracked now is how fast Huawei can manufacture enough chips to meet the demand DeepSeek has created, a question Relve, an AI tools and trends intelligence platform, is monitoring as China’s domestic AI stack accelerates.
