Chinese AI models continue narrowing the gap with the U.S. in cost efficiency and token consumption
AI summary card
Chinese AI models continue narrowing the gap with the U.S. in cost efficiency and token consumption
Jefferies tracks data from OpenRouter, Artificial Analysis, QuestMobile, Omdia, and F&S, and believes high-cost-performance Chinese large models, AI Agents, and coding scenarios are driving token consumption and cloud MaaS opportunities.
- OpenRouter data shows that total token consumption in the week of June 15 rose 4.7% week over week to 46.7tn; Chinese models reached 18.8tn, higher than 5.8tn for U.S. models.
- The Silicon Data LLM Token Expenditure Index stayed in the 1.64 to 1.68 range from June 14 to 19, below 2.04 on May 31, indicating a softer short-term spending mix.
- The report emphasizes that Chinese models such as DeepSeek, Qwen, Kimi, MiniMax, and Zhipu have clear cost-performance advantages in API cost relative to U.S. models, thanks to optimizations in MoE, attention mechanisms, and MFU.
- Jefferies focuses on the Volcano Engine FORCE conference, Doubao daily token consumption, Seedance 2.0, TRAE/Arkclaw/coze, as well as Tencent Weixin AI testing and a potential formal launch in 4Q26.
Report interpretation
Overview
This report is the 51st piece in Jefferies' AI Series, focusing on trends in cost performance, token consumption, and model intelligence, with a comparison between Chinese and U.S. AI models. The report integrates more than 50 data points, covering OpenRouter token consumption, API pricing across models, model intelligence and cost efficiency from Artificial Analysis, user trends in AI mobile applications, video generation models, AI coding and Agent products, as well as cloud MaaS and the outlook for the Chinese large-model market.
Core views
The core view is that AI use cases are expanding from "chat" to "work," and coding, Agents, Workspace, and medium- to long-form content generation will continue to drive token consumption growth; Chinese large models are continuing to catch up with U.S. models in intelligence while maintaining significant API cost advantages; cloud service providers and large-model platforms will benefit from rising model call volumes, but the weakening short-term token expenditure index, internal enterprise token usage caps, and price competition suggest that commercialization quality still needs to be verified.
Analysis framework
The report uses a cross-validation approach based on multi-source data: OpenRouter is used to observe token consumption and model share, the Silicon Data index is used to track changes in spending mix, Artificial Analysis and the Stanford AI Index are used to compare model intelligence and cost efficiency, QuestMobile/aicpb.com is used to track user trends in AI applications and internet applications, and forecasts from Omdia, Grand View Research, F&S, CAICT, and IDC are combined to assess the medium- to long-term market opportunity for Agentic AI, AI coding tools, MaaS, and the Chinese large-model market.
Methodology notes
Measure real demand for large-model calls through weekly token consumption, model rankings, and company share.
The report uses OpenRouter data for the week of June 15 to compare consumption between Chinese and U.S. models, and identifies high-usage models such as DeepSeek V4 Flash, MiMo V2.5, and MiniMax M3.
Compare model intelligence, input/output API pricing, and cost efficiency simultaneously.
The report believes that Chinese models such as DeepSeek, Qwen, Kimi, MiniMax, and Zhipu achieve lower API costs relative to U.S. models through architectural optimizations such as MoE, Group-Query Attention, sparse attention, linear attention, and MFU.
Use DAU, MAU, month-over-month, and year-over-year changes to observe user activity in AI applications and internet sub-sectors.
The report tracks AI applications such as Doubao, Qwen, Yuanbao, and DeepSeek, as well as monthly user changes in e-commerce, music, video, short-video, travel, feed, and search applications.
Use third-party forecasts to assess the medium- to long-term opportunity for Agentic AI, AI coding tools, MaaS, and the Chinese large-model market.
The report cites growth forecasts for the Agentic AI software market, the global AI coding tools market, the Chinese AI code generation market, and the Chinese LLM market to support its view on commercialization opportunities in cloud and enterprise AI.
Asset mapping & comparison
Structured mapping from thesis to named assets (strengths, weaknesses, peers, risks).
- Baidu, Alibaba, Tencent (BAT)The report believes BAT are beneficiaries of growth in AI model and Agent token consumption.
- Strengths
- They have cloud computing, models, application entry points, and ecosystem resources, enabling them to capture MaaS, Agent, and enterprise AI demand.
- Weaknesses
- Some application user metrics are diverging; for example, Baidu app MAU fell 8% YoY and DAU fell 21% YoY in May.
- Comparison
- Compared with standalone model companies, BAT have a more complete closed loop across cloud, models, and applications.
- Risks
- Soft token expenditure, enterprise budget constraints, model price competition, and slower-than-expected commercialization.
- Alibaba, Qwen, Alibaba CloudThe report highlights Alibaba's advantages in full-stack AI capabilities, enterprise token share, and Agentic AI platforms.
- Strengths
- It has a full-stack layout spanning chips and infrastructure, Qwen models, and applications; F&S shows Alibaba Cloud's enterprise token share reached 32% in 2H25; Omdia shows it ranked highest in 5 Agentic AI indicators.
- Weaknesses
- Consumer willingness to pay remains low, and monetization of enterprise customization and cloud revenue still needs time.
- Comparison
- In China's enterprise MaaS and Agent platform dimensions, the report views Alibaba Cloud as one of the leading participants.
- Risks
- Capital expenditure, overseas expansion execution, declining model prices, and weaker-than-expected enterprise AI project rollout.
- Tencent, Weixin AIWeixin AI testing is viewed as an important step for Tencent in advancing A2A and AI entry points.
- Strengths
- Weixin has high-frequency social and communication scenarios and can expand into search, productivity, entertainment, shopping, and tool generation.
- Weaknesses
- It is currently still in a small-scale testing stage, with formal release expected in 4Q26, and the commercialization path has not yet been fully disclosed.
- Comparison
- Compared with standalone AI assistants, Weixin AI has the advantage of a super-app entry point.
- Risks
- Launch delays, user experience below expectations, model capability, and privacy compliance requirements.
- Volcano Engine, Doubao, Seedance, TRAE, Arkclaw, cozeThe report treats the Volcano Engine FORCE conference as a key near-term observation point for AI product and commercialization updates.
- Strengths
- Doubao disclosed daily token consumption of 120tn in March 2026, and Doubao DAU grew 5% month over month to 158m in May.
- Weaknesses
- MaaS revenue targets, subscription package pricing, and video generation ARR trends still require further disclosure.
- Comparison
- It has a standout traffic advantage in application activity and model call volume.
- Risks
- Unclear pricing strategy, intensifying competition in video generation, token cost control, and weaker-than-expected enterprise customer conversion.
- Chinese large models such as DeepSeek, Kimi, MiniMax, and ZhipuThe report believes Chinese models are a key industry variable in narrowing the intelligence gap and improving cost efficiency.
- Strengths
- DeepSeek V4 Flash ranked first in OpenRouter model token consumption in the week of June 15; Chinese models achieve low API costs through optimizations such as MoE, attention mechanisms, and MFU.
- Weaknesses
- The report still notes that Chinese models trail U.S. models by about 7 months, and the Stanford AI Index shows top U.S. models leading Chinese models by 2.7%.
- Comparison
- Relative to U.S. models, Chinese models have stronger API cost advantages, but top-tier intelligence levels still require continued catch-up.
- Risks
- Model homogenization, price wars, inference cost pressure, and continued iteration by top overseas models.
- Kingsoft Cloud, AI labs, KlingThe report lists them as potential beneficiaries of rising token consumption and stronger demand for AI models/Agents.
- Strengths
- Cloud services, AI labs, and video generation models can benefit from the expansion of enterprise and content generation scenarios.
- Weaknesses
- Revenue realization depends on industry demand, customer budgets, and product differentiation.
- Comparison
- Compared with BAT, these assets may be more focused on specific cloud services or vertical AI scenarios.
- Risks
- Market competition, cost investment, customer concentration, and insufficient paid conversion for AI applications.
Key data
- Silicon Data LLM Token Expenditure Index1.64-1.68Remained soft from June 14 to 19, below 2.04 on May 31.
- OpenRouter weekly total token consumption46.7tn, +4.7% WoWGrowth in the week of June 15 relative to the week of June 8.
- Token consumption of Chinese models vs. U.S. modelsChina 18.8tn, United States 5.8tnFrom June 15 to 21, token consumption of Chinese models increased 2.12% week over week.
- OpenRouter company shareDeepSeek 18.5%, Anthropic 14.1%, Xiaomi 9.6%, Google 8.9%, MiniMax 8.7%DeepSeek ranked first by company-level token consumption share.
- OpenRouter model rankingDeepSeek V4 Flash 4.94tn, MiMo V2.5 3.94tn, MiniMax M3 3.77tnFollowed by Hy3 preview 3.63tn and Owl Alpha 2.56tn.
- Doubao large model daily token consumption120tnDaily token consumption level disclosed by Volcano Engine in March 2026.
- AI assistant DAUDoubao 158m, DeepSeek about 30.5m, Qwen 27.9m, Yuanbao 8.8mData for May 2026; Doubao +5% MoM, DeepSeek about +6% MoM.
- Agentic AI software marketUSD271m in 2025 to USD9.7bn in 2030, CAGR 105%Omdia forecast; Alibaba Cloud ranked highest in 5 out of 7 indicators.
- Global AI coding tools marketUSD26bn in 2030, about 3x 2025Grand View Research forecast; F&S expects China's AI code generation market to reach about USD4.7bn in 2028.
- China LLM market2024-2030 CAGR about 64%, exceeding RMB100bn by 2030F&S/CAICT forecast; enterprise-side growth is expected to outpace consumer-side growth.
- Enterprise daily token consumption37tn, +263% vs. half-year earlierF&S data; Alibaba Cloud's share rose to 32% in 2H25, ranking first.
Impact & implications
If AI usage expands from chat into coding, Agents, and enterprise workflows, growth in token consumption will increase revenue elasticity for cloud service providers, model platforms, and AI application companies. The high cost efficiency of Chinese models helps expand call volume, lower barriers to enterprise deployment, and strengthen the competitiveness of domestic cloud vendors in MaaS and enterprise AI. However, the weakening expenditure index and enterprise token caps suggest that short-term commercialization may still be constrained by budgets, and it remains necessary to observe whether high call volumes can convert into sustainable revenue and profit.
Risks
- The Silicon Data LLM Token Expenditure Index has weakened since the end of May, indicating that the short-term spending mix may be less optimistic than call volume suggests.
- Media reports indicate that some technology and internet companies have set caps on employee token consumption, which may suppress internal enterprise usage intensity.
- Lower API costs for Chinese models support adoption, but may also bring price competition and margin pressure.
- Consumer willingness to pay for AI applications remains low, and there is uncertainty around revenue recognition and delivery cycles for enterprise customized solutions.
- Although the intelligence gap between Chinese models and top U.S. models is narrowing, it may still be affected by rapid iteration of overseas models.
- If cloud vendors' MaaS revenue targets, video generation ARR, and Agent commercialization progress fall short of expectations, the investment narrative will weaken.
What to watch
- Volcano Engine FORCE conference on June 23-24: Doubao daily token consumption, subscription pricing, Seedance 2.0 ARR, TRAE/Arkclaw/coze, and MaaS revenue targets.
- Progress of Tencent Weixin AI small-scale testing, and the pace of formal launch in 4Q26.
- Changes in OpenRouter token consumption share between Chinese and U.S. models, top-model rankings, and the proportion of coding/Agent scenarios.
- Whether the Silicon Data LLM Token Expenditure Index rebounds from the 1.64-1.68 range.
- Monthly DAU/MAU trends of AI applications such as Doubao, Qwen, Yuanbao, and DeepSeek.
- Alibaba Cloud's delivery on enterprise token share, Agentic AI platforms, and the launch of Agentic services in Europe.
- Whether market size forecasts for AI coding tools, Agentic AI software, the China LLM market, and the MaaS market are validated by real revenue.