China’s LLM competition is entering a phase that equally emphasizes scale, Agent integration, and the cost curve
AI summary card
China’s LLM competition is entering a phase that equally emphasizes scale, Agent integration, and the cost curve
Nomura’s expert call argues that the next stage of competition in China’s large language models will be determined not only by parameter scale, but also by Agent toolchains, real workflow data flywheels, and system-level inference cost optimization.
- Chinese LLMs may move toward architectures with multiple trillions of parameters, with around 3 trillion parameters viewed by the expert as a practical milestone for the next stage.
- The marginal returns of simply expanding parameter scale are declining, while architectural efficiency, post-training, memory management, reliable tool use, and long-task capabilities are becoming more important.
- DeepSeek is viewed as having a leading price-performance advantage, with the expert estimating its blended inference cost at below RMB 0.8 per million tokens and its blended price at about RMB 1.3 per million tokens.
- Agent workflows can accumulate task traces, tool calls, user corrections, and failure cases, helping models and products build proprietary data flywheels.
- DeepSeek’s overseas growth is mainly developer-driven, with overseas markets contributing about 47% to 48% of ARR, and it may put pricing pressure on overseas closed-source models in execution-layer scenarios with high token consumption.
Report interpretation
Overview
This report summarizes key points from a call between Nomura’s China Internet team and an expert from a leading Chinese large language model company. The expert believes competition in China’s LLM market will revolve around three main themes: continued expansion of model and compute scale, deeper integration of models and Agent toolchains, and stronger system-level token economics. Although DeepSeek is unlisted, the expert views it as a representative player at the intersection of these trends, aiming to narrow the capability gap with leading overseas models while maintaining structurally low costs and lower token pricing.
Core views
The core view is that future competitive advantage in LLMs will shift from standalone model capability to integrated system capability spanning models, Agent toolchains, real user interactions, and the cost curve. Parameter scale remains important, but marginal returns are declining; more reliable tool use, long-horizon task execution, automated evaluation, and extraction of high-quality training signals will determine iteration speed. Through optimization of attention and memory management, offloading KV-cache to NvMe SSDs, improving infrastructure utilization, and using domestic accelerators, DeepSeek may be able to convert cost advantages into pricing advantages and developer penetration.
Analysis framework
The report uses expert interviews and industry trend synthesis to analyze China’s LLM ecosystem across five dimensions: model scale, product form, cost structure, overseas adoption, and open-source strategy. Its conclusions are mainly based on the expert’s estimates of DeepSeek’s costs, ARR, user mix, model release cadence, and developer workflows.
Methodology notes
Declining marginal returns from parameter scale
The report argues that models will continue to expand to multiple trillions of parameters, but the returns from simply piling on parameters are declining, and competition will depend more on architectural efficiency, post-training efficiency, memory management, and long-task performance.
Agent workflow data flywheel
Once models are embedded into Agent toolchains, task traces, tool calls, user corrections, execution failures, and real workflows can feed back into training and product optimization, shifting the advantage from the base model to the integrated system.
Inference cost and token pricing
Using cache hit rates, inference costs, and pricing, the expert estimates DeepSeek’s contribution margin, showing that low-cost inference capability may support its pricing leadership.
Open source lowers adoption barriers
An open-source strategy may compress direct monetization potential, but it can expand overseas reach, support multi-cloud and multi-hardware deployment, and weaken the ability of closed-source models to sustain high API premiums.
Asset mapping & comparison
Structured mapping from thesis to named assets (strengths, weaknesses, peers, risks).
- DeepSeekUnlisted LLM company and the core subject of discussion in the report
- Strengths
- Low cost curve, competitive token pricing, strong developer adoption, open-source strategy favorable for distribution, and the ability to build a data flywheel through Agent workflows.
- Weaknesses
- Penetration into large enterprises remains relatively limited; the delayed V4 release indicates the Agent toolchain still requires debugging and optimization; direct monetization may be constrained by the open-source strategy.
- Comparison
- Compared with overseas products such as Claude Code or Codex, DeepSeek may first become complementary in token-intensive areas such as sustained code generation and execution, rather than replacing leading overseas models across the full workflow.
- Risks
- Capability gap narrowing may fall short of expectations, deployment efficiency of domestic accelerators may disappoint, price declines may compress profits, and overseas regulation or customer data requirements may affect adoption.
- Overseas closed-source frontier model providersDeepSeek’s benchmark and potential target of pricing pressure
- Strengths
- They still retain premium pricing defensiveness in complex reasoning, planning, architecture design, and highly reliable tasks.
- Weaknesses
- In execution-layer and high-token-consumption tasks that can be handled by lower-cost models, API premiums may come under pressure.
- Comparison
- Developers may first use high-end overseas models for planning and complex reasoning, then switch to DeepSeek for sustained code generation and execution.
- Risks
- As task routing becomes more widespread, token share for high-quality but expensive models may be diverted to lower-cost models.
- China Internet & New Media sectorIndustry covered by this report
- Strengths
- It has large-scale user scenarios, application data, and a developer ecosystem, which are conducive to model product iteration and Agent deployment.
- Weaknesses
- Model capability, compute clusters, infrastructure interconnectivity, and product reliability still require sustained investment.
- Comparison
- Chinese LLM players are attempting to compete in a differentiated way against leading overseas models through cost efficiency and open-source diffusion.
- Risks
- Declining marginal returns from pure parameter expansion, commercialization price wars, shorter technology leadership cycles, and regulatory and cross-border data requirements.
Key data
- Report date2026-07-27The report cover page shows the date as 27 July 2026.
- Next-stage parameter scale milestoneAbout 3 trillion parametersThe expert believes about 3tn parameters is a practical milestone for the next stage of Chinese LLMs.
- DeepSeek blended inference costBelow RMB 0.8 per million tokensThe estimate assumes a 60% cache hit rate and excludes model training compute, training-related personnel expenses, and other R&D costs.
- DeepSeek blended priceAbout RMB 1.3 per million tokensBased on this, the expert estimates an inference contribution gross margin of about 40%.
- V4 generation price declineAbout 75%The expert says that even if V4 generation pricing declines by about 75%, DeepSeek could still maintain an inference contribution gross margin of about 40%.
- DeepSeek ARRAbout USD 520 millionThe expert estimates DeepSeek’s ARR at about USD520mn as of June 2026.
- Overseas market ARR contributionAbout 47% to 48%Overseas users are concentrated in Europe and the United States, followed by Japan, South Korea, and Australia.
- Leadership duration of a single model generationShortened from about 6 months to 3 to 4 monthsAI-in-the-loop, automated evaluation, LLM-as-a-Judge, multi-model scoring, and teacher-model feedback are accelerating model iteration.
Impact & implications
The investment implication is that competitive assessment of AI model vendors should shift from pure capability leaderboards to system-level commercialization capability: who can handle more token-intensive tasks at lower cost, who can build proprietary data through Agent products, and who can create distribution advantages within the developer ecosystem. Low-cost open-source models may not fully replace leading overseas models, but they can gain share in code generation, execution, and other high-frequency token-consumption scenarios, while putting pressure on closed-source model API pricing.
Risks
- The report is mainly based on expert interviews, and DeepSeek’s costs, ARR, and user mix are expert estimates that may not be the same as company-disclosed data.
- A low-price strategy may expand adoption, but it could also depress industry profit margins and direct monetization potential.
- An open-source strategy helps distribution and ecosystem expansion, but it may weaken API monetization and closed-source premiums.
- If Agent toolchains lack sufficient debugging and reliability, deployment in professional software development and knowledge-work scenarios may be affected.
- The shortening of model leadership cycles may make it difficult for the generational advantage of a single model to be sustained over the long term.
- Overseas adoption may be affected by data control, compliance, geopolitics, and enterprise procurement processes.
What to watch
- DeepSeek V4’s formal release cadence, Agent toolchain stability, and product experience for professional developers.
- Whether Chinese LLMs move toward about 3 trillion parameters as the expert expects, along with corresponding compute clusters, supernodes, high-speed interconnects, and optical network construction.
- Changes in DeepSeek’s actual token pricing, cache hit rates, inference costs, and contribution gross margins.
- Whether overseas developer adoption expands from individuals and small teams to large enterprises.
- Whether developers form a layered usage pattern in which high-end overseas models handle planning while DeepSeek handles execution.
- Whether pressure from open-source models on closed-source model API pricing and token share intensifies.
- The actual effectiveness of AI-in-the-loop, automated evaluation, LLM-as-a-Judge, and teacher-model feedback in model iteration.