Bernstein updates China AI inference token economics model, with conclusions still favoring frontier models and high-throughput optimization capabilities
AI summary card
Bernstein updates China AI inference token economics model, with conclusions still favoring frontier models and high-throughput optimization capabilities
The report incorporates feedback from leading Chinese AI labs into its inference margin model, arguing that Minimax's 40% margin requires materially higher TPS assumptions, while Z.ai may still achieve higher margins under comparable discounts, and that DeepSeek V4 and Tencent hy3-preview highlight continued acceleration in China's AI competition.
- For Minimax M2.7 to reach about 40% gross margin in text/code inference, it would require a data-center-level TPS ratio of roughly 4.5-5.0x versus GLM-5/5.1.
- The report adds a commercial discount variable, with a base-case assumption of a uniform 20% discount; higher discounts would directly reduce revenue per million tokens and inference margins.
- Under comparable discount, workload, and TPS assumptions, Bernstein still believes Z.ai's GLM series may have higher margins than Minimax, though the gap has narrowed versus previous estimates.
- DeepSeek V4 is close to the global SOTA of several months ago while costing less, and the low pricing of V4-Flash reinforces the view of a price war in lightweight low-cost models.
- Tencent hy3-preview is seen as a step in the right direction, albeit slightly behind; the key issue to watch is whether Tencent's AI data and infrastructure rebuild can enable faster model iteration.
Report interpretation
Overview
This report is Bernstein's in-depth supplementary study on China Internet and AI inference economics, with the core focus of updating the AI token economics and inference margin model based on feedback from leading Chinese AI labs. The report discusses Minimax, Z.ai, DeepSeek V4, and Tencent hy3-preview, and evaluates inference margins and competitive implications across different models by incorporating variables such as commercial discounts, TPS, GPU costs, KV cache, and input/output token pricing.
Core views
The core views include: first, management's claim of about 40% inference gross margin for Minimax M2.7 requires higher tokens-per-second assumptions than previously modeled; second, Z.ai's GLM series may still achieve higher margins under comparable discounts and workloads, though the relative advantage is smaller than previously estimated; third, multimodal inference usually has higher margins than text/code inference because revenue per million tokens is higher; fourth, DeepSeek V4 shows that Chinese model capabilities remain close to the global frontier, but price competition among lightweight low-cost models will become more intense; fifth, the investment implication of Tencent hy3-preview lies not in the single release itself, but in whether it can prove faster iteration of Tencent's internal AI models and applications.
Analysis framework
The report uses a bottom-up AI token economics model, combining variables such as revenue per million tokens, input/output pricing, KV cache hit rates, commercial discounts, TPS, GPU utilization, and GPU-hour costs to estimate inference costs and gross margins for different models, and then uses sensitivity analysis to observe the effect of TPS and discount changes on margins. At the same time, the report calibrates assumptions using Lambda.ai throughput benchmarks, company feedback, and model pricing comparisons.
Methodology notes
Revenue per million tokens and inference cost model
Uses input/output token pricing, input-output length ratios, KV cache discounts, commercial discounts, TPS, GPU utilization, and GPU-hour costs to estimate revenue, cost, and inference gross margin per million tokens.
TPS and commercial discount sensitivity analysis
Adjusts tokens per second and blended commercial discounts to observe implied margin changes for Minimax M2.5/M2.7 and Z.ai GLM-5/5.1.
Competition between frontier intelligence and low-cost lightweight models
The report distinguishes between frontier high-capability models and lightweight low-cost models, arguing that frontier intelligence remains the key long-term determinant of winners and losers, while price competition among lightweight models will intensify.
Asset mapping & comparison
Structured mapping from thesis to named assets (strengths, weaknesses, peers, risks).
- Tencent Holdings Ltd / 700.HKCovered company and observation target for AI application/model iteration
- Strengths
- Rated Outperform with target price HK$780; feedback on AI applications in advertising is positive, and hy3-preview shows improvement in the direction of model iteration.
- Weaknesses
- hy3-preview looks more like a transitional product, and the market is still waiting for stronger models and progress in WeChat agentic capabilities.
- Comparison
- Compared with domestic AI peers such as DeepSeek, Z.ai, and Kimi, Tencent still needs to prove the iteration speed of its in-house models.
- Risks
- Macro consumption, user engagement, competition in gaming and advertising, and regulatory risks such as antitrust.
- Alibaba / BABA / 9988.HKCovered company and relevant play in China's internet AI ecosystem
- Strengths
- Rated Outperform with target price US$180/HK$176; the cloud business and AI investment portfolio support valuation.
- Weaknesses
- Losses in innovative businesses and competition in core e-commerce still affect earnings visibility.
- Comparison
- The report mentions Alibaba's investments in Kimi and Minimax, contrasting with Tencent potentially leading DeepSeek financing in terms of ecosystem positioning.
- Risks
- Macro consumption, platform user engagement, competition, regulation, and losses in innovative businesses.
- Z.ai GLM-5/GLM-5.1Comparison target for AI lab model margins
- Strengths
- Under comparable discount and TPS assumptions, the report believes the GLM series may achieve relatively high inference margins.
- Weaknesses
- Margins are highly sensitive to TPS, commercial discounts, and workload mix.
- Comparison
- Versus Minimax, the margin advantage remains but has narrowed relative to previous estimates.
- Risks
- Price competition, expanding discounts, and throughput assumptions falling short of expectations.
- Minimax M2.5/M2.7Calibration target for AI inference margin models
- Strengths
- Company feedback indicates M2.7 text/code inference gross margin can reach about 40%, supported by relatively high TPS assumptions.
- Weaknesses
- Achieving a 40% margin requires a high TPS ratio, and the model is sensitive to discount and workload changes.
- Comparison
- M2.5's single-GPU Lambda benchmark throughput is about 5.1x that of GLM-5, yet Z.ai model margins may still be higher.
- Risks
- Expanded commercial discounts and a higher share of text/code workloads causing the revenue mix to be less favorable than multimodal.
- DeepSeek V4Observation target for China's AI frontier capability and price competition
- Strengths
- Close to the global SOTA of several months ago while costing less; open source/open weights help lift the overall capability of Chinese AI labs.
- Weaknesses
- The capability gap versus domestic peers is narrowing, and low-price strategy may weaken pricing power.
- Comparison
- V4-Pro pricing is close to domestic leading models such as GLM-5.1 and Qwen3.6 Max; V4-Flash is near the low end of lightweight low-cost model pricing.
- Risks
- Low prices and limited-time discounts intensify industry competition, and long-term profitability remains to be proven.
Key data
- Minimax M2.7 base-case inference margin42.5%Based on model assumptions of a 20% commercial discount, TPS of 9,000, GPU utilization of 90%, and GPU cost of $1.6 per hour.
- Z.ai GLM-5.1 base-case inference margin50.3%The report believes the GLM series may still achieve higher margins than Minimax models under comparable discount assumptions.
- Z.ai GLM-5 base-case inference margin33.9%Based on the report's updated token economics model.
- Minimax M2.5 base-case inference margin29.5%Base case derived after the report raised its H1 2026 TPS assumption.
- GPU cost assumption$1.6/GPU-hourCompany feedback suggests this estimate is broadly reasonable, and the report assumes leading Chinese AI labs have broadly similar GPU costs.
- Base commercial discount assumption20%The report adds a commercial discount variable and notes that customer mix and limited-time promotions can materially affect revenue per token.
- Lambda.ai throughput benchmarkM2.5 about 4,031 tok/s/GPU; GLM-5 about 788 tok/s/GPUAccording to the report, M2.5 single-GPU throughput is about 5.1x that of GLM-5.
- Tencent target priceHK$780Based on 20x FY+1 PE valuation; rated Outperform.
- Alibaba target priceUS$180/HK$176Based on SOTP valuation of FY+1 revenue and profit for the core e-commerce and cloud businesses; rated Outperform.
Impact & implications
The investment implication is that Chinese AI model capabilities and commercialization competition are still evolving rapidly, and inference margins cannot be judged by list prices alone, but must also take into account TPS, discounts, workloads, and GPU utilization. For internet leaders, frontier model capabilities, the speed of AI infrastructure iteration, application-layer distribution entry points, and commercialization efficiency will shape the medium- to long-term valuation narrative; however, price wars in low-cost models, commercial discounts, and regulatory/macro risks will weigh on earnings visibility.
Risks
- Intensifying price competition among low-cost lightweight AI models, reducing revenue per million tokens.
- Commercial discounts, customer mix, and limited-time promotions may cause external margin estimates to diverge from actual financial results.
- Key assumptions such as TPS, GPU utilization, and KV cache hit rates are difficult to verify precisely from the outside.
- China's internet leaders still face risks from macro consumption, regulation, platform competition, and fluctuations in user engagement.
- How AI lab costs are classified among cost of sales, R&D expense, and marketing expense may affect comparability.
What to watch
- Actual inference margins and discount levels disclosed later by AI labs such as Minimax and Z.ai.
- Follow-up adoption of DeepSeek V4, pricing strategy, and potential financing arrangements.
- The release cadence of subsequent official versions and larger models after Tencent hy3-preview.
- Whether Tencent's reorganization of AI data and infrastructure leads to faster model iteration.
- Changes in the workload mix of multimodal, video generation, text-to-speech, and text/code tasks.
- Commercialization progress of AI among Chinese internet companies in advertising, gaming, the WeChat ecosystem, and cloud businesses.