China's large model competition has entered a parallel phase of scale, agent integration, and cost optimization
AI summary card
China's large model competition has entered a parallel phase of scale, agent integration, and cost optimization
Nomura's expert call believes that the next round of competition among Chinese LLM vendors will be driven jointly by multi-trillion-parameter scaling, deep integration of models with agent toolchains, and system-level token economics, with DeepSeek occupying a standout position in low cost and developer adoption.
- Chinese LLMs may move toward a scale of around 3tn parameters, but the marginal returns from simply expanding parameters are declining.
- The focus of model competition is shifting from standalone model capability to integrated systems combining models, tool calling, agent execution, and feedback from real workflows.
- Experts estimate DeepSeek's ARR at around USD520mn in June 2026, with overseas markets contributing about 47-48%.
- DeepSeek's blended inference cost is estimated at below CNY0.8 per million tokens, with a blended price of about CNY1.3 per million tokens and an inference contribution gross margin of about 40%.
- Low-cost open-source models may gain higher usage in token-intensive execution layers and put pressure on the API premium of overseas closed-source models.
Report interpretation
Overview
This report summarizes a call between Nomura's China Internet team and an expert from a leading Chinese LLM vendor. The core view is that competition in Chinese large models is shifting from a pure race in model parameters to a broader contest in compute scale, integration of models with agent toolchains, system-level inference cost optimization, and flywheels built from real user interaction data. DeepSeek is seen as a representative company at the intersection of these trends, and its low-cost architecture, developer-oriented overseas growth, and open-source strategy may change global LLM pricing and usage structures.
Core views
The report argues that Chinese LLMs will continue to scale up model size, with around 3tn parameters likely becoming the next practical milestone, but the marginal benefits of parameter scale are declining. Long-term competitiveness will depend more on post-training efficiency, memory management, reliable tool calling, long-horizon task performance, and integrated model-agent systems. The delayed formal release of DeepSeek V4 partly reflects the need for further debugging and optimization at the agent-harness layer, with the goal of building an integrated product similar to Claude Code or Codex that can cover professional software development and broader knowledge work tasks.
Analysis framework
The report uses expert interviews and industry trend analysis to assess the evolution of China's LLM industry from six angles: model scale, toolchain integration, cost structure, overseas adoption, open-source strategy, and changes in the competitive cycle. The analytical focus is not traditional financial valuation, but rather judging the competitive positioning of DeepSeek and Chinese LLM vendors through metrics such as inference cost, pricing, ARR, overseas revenue share, and user composition.
Methodology notes
Equal emphasis on parameter scaling and efficiency optimization
The report believes that multi-trillion-parameter models remain the direction of development for Chinese LLMs, but competitiveness will increasingly depend on architectural efficiency, post-training efficiency, memory management, and long-task execution capability.
model-harness integration
When models are combined with tool calling, task trajectories, user corrections, and feedback from execution failures, they can form a proprietary data flywheel, shifting competitive advantage from standalone models to integrated model-agent systems.
Inference cost, pricing, and contribution gross margin
The report uses inference cost and pricing per million tokens to measure price competitiveness, and notes that DeepSeek lowers its cost curve through caching, attention and memory management, NVMe SSD KV-cache offloading, and domestic accelerator deployment.
Asset mapping & comparison
Structured mapping from thesis to named assets (strengths, weaknesses, peers, risks).
- DeepSeekCore subject of the report discussion, unlisted
- Strengths
- It has price-performance advantages, lower inference costs, developer-oriented overseas growth, open-source distribution capability, and the ability to reduce the cost curve through architecture and infrastructure optimization.
- Weaknesses
- Penetration into large enterprises remains relatively limited; the delayed release of V4 shows that the agent-harness layer still requires debugging and optimization; direct commercialization may be constrained by the open-source strategy.
- Comparison
- Compared with high-end overseas model tools such as Claude Code and Codex, DeepSeek may be more focused on token-intensive areas such as continuous code generation and execution rather than fully replacing complex planning and high-reliability reasoning tasks.
- Risks
- Model capability gap, product integration progress, overseas compliance and customer trust, price competition, and training and R&D costs not included in the inference gross margin estimate.
- 海外 frontier modelsDeepSeek's benchmark and complementary peers
- Strengths
- They still retain premium room in complex reasoning, reliability, and high-quality task completion.
- Weaknesses
- High API pricing may come under pressure after low-cost open-source models become widespread.
- Comparison
- The report describes developers as possibly using Claude Code or Codex first for planning, architecture design, and complex reasoning, and then switching to DeepSeek to execute high-token-consumption code generation and implementation tasks.
- Risks
- If users become more accustomed to routing models by task within workflows, closed-source models' token share in the execution layer may be eroded by low-cost models.
- 中国 LLM 开发商Beneficiaries of industry trends
- Strengths
- They may benefit from larger compute clusters, supernodes, high-speed interconnects, optical networks, and domestic accelerator deployment.
- Weaknesses
- Marginal returns from parameter scaling are declining, the leadership window of a single model generation is shortening, and the competitive cycle is accelerating.
- Comparison
- Competitive advantage will shift from standalone model capability to integrated model-agent systems, automated evaluation, and real workflow training signals.
- Risks
- Compute supply, model debugging, cost control, commercialization paths, and the impact of open source on revenue conversion.
Key data
- Next-stage parameter scale milestone for Chinese LLMsAbout 3tn parametersExperts believe this is the next practical milestone for Chinese LLMs.
- Estimated ARR of DeepSeekAbout USD520mnExpert estimate as of June 2026.
- DeepSeek overseas market revenue contributionAbout 47-48%Overseas users are concentrated in Europe and the United States, followed by Japan, South Korea, and Australia.
- DeepSeek blended inference costBelow CNY0.8 per million tokensAssumes a 60% cache-hit rate and excludes model training compute, training-related personnel expenses, and other R&D costs.
- DeepSeek blended pricingAbout CNY1.3 per million tokensCorresponding to an estimated inference contribution gross margin of about 40%.
- DeepSeek V4 generational price reductionAbout 75%The report says that even with a significant price cut, DeepSeek can still maintain an inference contribution gross margin of about 40%.
- Leadership duration of a single model generationAbout 3-4 monthsExperts believe this has shortened from about 6 months, with AI-in-the-loop being one of the main drivers.
Impact & implications
The investment implication of the report is that value in the LLM industry may gradually shift from pure model parameters and closed-source API premiums toward cost curves, toolchain integration, real workflow data, and developer ecosystems. Low-cost, open-source, developer-oriented models such as DeepSeek may not fully replace overseas frontier models, but they may capture substantial share in execution layers with higher token consumption and stronger price sensitivity, thereby pushing down global inference service prices and reshaping the layered use of models.
Risks
- The marginal returns from simply expanding parameters are declining, which may weaken the competitive advantages brought by scaled training.
- DeepSeek's cost and gross margin estimates do not include model training compute, training-related personnel expenses, or other R&D costs, and therefore cannot be directly equated with the company's full gross margin.
- Low-price and open-source strategies may increase adoption, but they may also suppress direct commercialization revenue.
- Penetration into large enterprise customers remains limited, and overseas growth is currently driven more by developers and small to mid-sized project teams.
- The model leadership cycle has shortened to about 3-4 months, and faster industry competition may increase ongoing R&D pressure.
What to watch
- DeepSeek V4's formal release timing, agent-harness stability, and product form.
- Growth in DeepSeek's overseas ARR, changes in overseas revenue share, and penetration into large enterprise customers.
- Inference cost per million tokens, pricing, and the pace of price adjustments.
- Pressure from open-source models on closed-source API premiums and global model pricing.
- The speed of application of AI-in-the-loop, LLM-as-a-Judge, multi-model scoring, and teacher-model feedback in model iteration.
- The impact of domestic accelerators, NVMe SSD KV-cache offloading, optical networks, and high-utilization infrastructure on the cost curve.