Introduction
The domestic large language model industry entered a new phase of price competition in 2026, triggered by sweeping price cuts across major model service providers. At a recent investment summit, Liang Wenfeng, founder of DeepSeek, highlighted sky-high inference chip costs as a critical industry pain point. Media reports also revealed that DeepSeek has secretly launched its custom inference chip development program for roughly one year.
Inference silicon dominates operational expenditure for LLM operators. Gaining control over chip costs has evolved into one of the most decisive competitive advantages for model vendors. This article analyzes the ongoing price competition wave driven by DeepSeek, unpacks the engineering optimizations behind low-price model services, explains why inference hardware restricts profit margins, and outlines the formidable obstacles that AI companies face when building self-developed chips.
1. DeepSeek Sparks the Industry-wide Price Reduction Wave
A broad round of price cuts swept China’s LLM market starting in May 2026. On May 23, DeepSeek rolled out permanent price reductions for its V4-Pro model alongside a peak-valley dynamic pricing mechanism. Soon after, competitors including Xiaomi and Tencent Cloud followed suit by lowering their API service rates.
After adjustment, DeepSeek’s professional edition pricing dropped to 0.025 CNY per million tokens, merely 5.8% of MiniMax M2.7’s cost level. Its V4-Flash variant also secured top rankings on global model invocation leaderboards after the pricing update.
The price reshuffle forced all market participants to rethink cost structures. Before this round of adjustments, model service pricing relied heavily on brand premium; today, computing resource expenditure directly determines pricing power. Many enterprise customers now route workloads across multiple LLM vendors to balance cost and performance. An API gateway such as 4sapi enables unified traffic scheduling across heterogeneous model endpoints, helping businesses flexibly switch providers amid frequent pricing fluctuations.
2. The Two Core Pillars Behind DeepSeek’s Low-Cost Service
DeepSeek’s ability to sustain aggressive low pricing stems from two coordinated strategies: algorithm-level optimization and domestic hardware supply chain adoption. First, continuous algorithm and engineering optimization slashes raw computing consumption. For its V4 product line processing million-token ultra-long context tasks, overall computing resource consumption falls to just 27% of its prior generation model. Optimizations cover attention mechanism refinement, KV cache compression, and dynamic batching strategies, which directly cut the hardware resources required for each inference request.
Second, the company expanded procurement of domestic heterogeneous AI accelerators such as Huawei Ascend chips to replace part of imported hardware. Localized hardware procurement lowers long-term capital expenditure, and also accelerates the commercial iteration of domestic AI infrastructure ecosystems.
These dual measures form a replicable blueprint for other model operators: algorithm improvements reduce theoretical resource demand, while domestic chips bring down hardware purchase expenses. However, this approach can only deliver limited gains, as the ceiling for software optimization gradually emerges. Further cost reduction requires breakthroughs at the hardware layer.
3. Inference Chips: The Primary Cost Barrier for LLM Operators
As algorithm convergence accelerates and domestic alternative hardware matures, gaps between model vendors’ algorithm advantages continue to narrow. Under this trend, inference chips have become the dominant component of operating costs.
Two key trends drive this shift. For cloud service operators including Alibaba Cloud and Meta, budgets allocated to inference chip procurement keep rising alongside growing user request volumes. More critically, throughout the complete lifecycle of model services, inference hardware accounts for 80% to 90% of total operational expenses within production environments.
Unlike training clusters, inference workloads run 24/7 to process user API requests. Continuous power consumption, hardware depreciation, and maintenance costs create persistent financial pressure. Even with advanced algorithm tuning, model operators cannot escape hardware cost constraints in the long run. This explains why multiple leading LLM firms have begun exploring self-developed chip roadmaps.
4. The Daunting Obstacles Facing Custom Inference Chip Development
A growing number of AI companies such as DeepSeek and OpenAI are exploring self-designed inference silicon to break hardware cost shackles. Nevertheless, the path carries substantial, multi-dimensional challenges.
First, the project requires teams to expand capabilities from pure software development into hardware R&D. Companies must recruit silicon architects, verification engineers, and board design specialists; knowledge accumulation cycles for cross-domain talent are lengthy. Existing AI research teams focused on model training lack hardware engineering experience.
Second, semiconductor hardware follows extremely fast iteration cycles. LLM companies risk the situation where newly taped-out chips become technically outdated immediately after mass production. Unlike model iteration that can deploy incremental updates online, chip development spans multiple years, making it hard to keep pace with rapidly evolving LLM architectures.
Third, independent chip development lacks sufficient scale economic support. Compared with dedicated semiconductor firms, AI model vendors have limited internal chip consumption volume. Without external bulk sales revenue, unit manufacturing costs remain difficult to push down, potentially offsetting expected long-term cost savings.
All these risks explain why only top-tier model enterprises dare to invest heavily in silicon research, while most medium-sized LLM teams remain reliant on third-party off-the-shelf accelerators.
5. The Industry Enters an Era of Cost Competition
The year 2026 marks a turning point: the large model industry is shifting its core competition dimension from "building capable models" to "controlling end-to-end operating costs."
Whether a model vendor can successfully develop and deploy self-designed inference chips will define its corporate positioning. Enterprises that master both model algorithms and underlying hardware evolve into full-stack AI infrastructure providers. Companies confined only to model software services will remain pure AI product vendors, facing sustained margin pressure amid price wars.
The industry divergence presents two distinct development paths. Teams without capital reserves for chip R&D will focus on algorithm engineering optimization and domestic hardware adaptation to maximize utilization efficiency of existing accelerators. Well-funded leading firms will persist in silicon development to seize long-term cost advantages.
Conclusion
Inference chip costs have become a decisive bottleneck restricting the sustainable growth of large model businesses. While self-developed hardware brings enormous challenges in talent, capital, and technical iteration, it is increasingly viewed as a necessary strategic investment for top-tier players.
In the coming years, the gap between vendors with independent chip capabilities and those relying entirely on commercial off-the-shelf hardware will widen further. Enterprises that achieve breakthroughs in hardware-software co-optimization will occupy a dominant position within the increasingly competitive AI industry. Meanwhile, medium-sized developers will continue to leverage algorithm optimization and flexible multi-hardware deployment strategies to survive in the fierce price competition.




