Quick Summary
Covering the latest research from top Wall Street investment banks

Bernstein updates China AI inference token economics model, with conclusions still favoring frontier models and high-throughput optimization capabilities

Institution
Bernstein
Date
2026-04-27
Authors
Robin Zhu, Charles Gou, Min-Joo Kang
Company
China Internet
Ticker
700.HK; BABA; 9988.HK
Industry
Internet Content & Information; Artificial Intelligence
Rating
Tencent Outperform; Alibaba Outperform
NeutralLow confidenceThe report believes Chinese AI labs remain close to the global frontier in model capabilities, Z.ai model margins may be higher than Minimax's, but price competition among low-cost lightweight models will remain intense.
AuthorsRobin Zhu, Charles Gou, Min-Joo Kang
Target priceTencent HK$780; Alibaba US$180/HK$176
Asset classesEquity
Business segmentsAI inference and token economics model、Chinese AI labs、Internet platforms、Cloud and AI infrastructure、AI applications in advertising and gaming
Research firm divisions/subsidiariesBernstein(Other)

AI summary card

Bernstein updates China AI inference token economics model, with conclusions still favoring frontier models and high-throughput optimization capabilities

The report incorporates feedback from leading Chinese AI labs into its inference margin model, arguing that Minimax's 40% margin requires materially higher TPS assumptions, while Z.ai may still achieve higher margins under comparable discounts, and that DeepSeek V4 and Tencent hy3-preview highlight continued acceleration in China's AI competition.

Tencent 700.HK rated Outperform with target price HK$780; Alibaba BABA/9988.HK rated Outperform with target price US$180/HK$176.
China InternetAI inferencetoken economicsDeepSeek V4Tencent hy3-previewZ.aiMinimaxprice competition
  • For Minimax M2.7 to reach about 40% gross margin in text/code inference, it would require a data-center-level TPS ratio of roughly 4.5-5.0x versus GLM-5/5.1.
  • The report adds a commercial discount variable, with a base-case assumption of a uniform 20% discount; higher discounts would directly reduce revenue per million tokens and inference margins.
  • Under comparable discount, workload, and TPS assumptions, Bernstein still believes Z.ai's GLM series may have higher margins than Minimax, though the gap has narrowed versus previous estimates.
  • DeepSeek V4 is close to the global SOTA of several months ago while costing less, and the low pricing of V4-Flash reinforces the view of a price war in lightweight low-cost models.
  • Tencent hy3-preview is seen as a step in the right direction, albeit slightly behind; the key issue to watch is whether Tencent's AI data and infrastructure rebuild can enable faster model iteration.

Report interpretation

Overview

This report is Bernstein's in-depth supplementary study on China Internet and AI inference economics, with the core focus of updating the AI token economics and inference margin model based on feedback from leading Chinese AI labs. The report discusses Minimax, Z.ai, DeepSeek V4, and Tencent hy3-preview, and evaluates inference margins and competitive implications across different models by incorporating variables such as commercial discounts, TPS, GPU costs, KV cache, and input/output token pricing.

Core views

The core views include: first, management's claim of about 40% inference gross margin for Minimax M2.7 requires higher tokens-per-second assumptions than previously modeled; second, Z.ai's GLM series may still achieve higher margins under comparable discounts and workloads, though the relative advantage is smaller than previously estimated; third, multimodal inference usually has higher margins than text/code inference because revenue per million tokens is higher; fourth, DeepSeek V4 shows that Chinese model capabilities remain close to the global frontier, but price competition among lightweight low-cost models will become more intense; fifth, the investment implication of Tencent hy3-preview lies not in the single release itself, but in whether it can prove faster iteration of Tencent's internal AI models and applications.

Analysis framework

The report uses a bottom-up AI token economics model, combining variables such as revenue per million tokens, input/output pricing, KV cache hit rates, commercial discounts, TPS, GPU utilization, and GPU-hour costs to estimate inference costs and gross margins for different models, and then uses sensitivity analysis to observe the effect of TPS and discount changes on margins. At the same time, the report calibrates assumptions using Lambda.ai throughput benchmarks, company feedback, and model pricing comparisons.

Methodology notes

  • AI unit economicsAI token economics model

    Revenue per million tokens and inference cost model

    Uses input/output token pricing, input-output length ratios, KV cache discounts, commercial discounts, TPS, GPU utilization, and GPU-hour costs to estimate revenue, cost, and inference gross margin per million tokens.

  • Sensitivity analysisTPS and discount sensitivity

    TPS and commercial discount sensitivity analysis

    Adjusts tokens per second and blended commercial discounts to observe implied margin changes for Minimax M2.5/M2.7 and Z.ai GLM-5/5.1.

  • Competition & strategyFrontier intelligence versus low-cost model competition

    Competition between frontier intelligence and low-cost lightweight models

    The report distinguishes between frontier high-capability models and lightweight low-cost models, arguing that frontier intelligence remains the key long-term determinant of winners and losers, while price competition among lightweight models will intensify.

Asset mapping & comparison

Structured mapping from thesis to named assets (strengths, weaknesses, peers, risks).

  • Tencent Holdings Ltd / 700.HK
    Covered company and observation target for AI application/model iteration
    Strengths
    Rated Outperform with target price HK$780; feedback on AI applications in advertising is positive, and hy3-preview shows improvement in the direction of model iteration.
    Weaknesses
    hy3-preview looks more like a transitional product, and the market is still waiting for stronger models and progress in WeChat agentic capabilities.
    Comparison
    Compared with domestic AI peers such as DeepSeek, Z.ai, and Kimi, Tencent still needs to prove the iteration speed of its in-house models.
    Risks
    Macro consumption, user engagement, competition in gaming and advertising, and regulatory risks such as antitrust.
  • Alibaba / BABA / 9988.HK
    Covered company and relevant play in China's internet AI ecosystem
    Strengths
    Rated Outperform with target price US$180/HK$176; the cloud business and AI investment portfolio support valuation.
    Weaknesses
    Losses in innovative businesses and competition in core e-commerce still affect earnings visibility.
    Comparison
    The report mentions Alibaba's investments in Kimi and Minimax, contrasting with Tencent potentially leading DeepSeek financing in terms of ecosystem positioning.
    Risks
    Macro consumption, platform user engagement, competition, regulation, and losses in innovative businesses.
  • Z.ai GLM-5/GLM-5.1
    Comparison target for AI lab model margins
    Strengths
    Under comparable discount and TPS assumptions, the report believes the GLM series may achieve relatively high inference margins.
    Weaknesses
    Margins are highly sensitive to TPS, commercial discounts, and workload mix.
    Comparison
    Versus Minimax, the margin advantage remains but has narrowed relative to previous estimates.
    Risks
    Price competition, expanding discounts, and throughput assumptions falling short of expectations.
  • Minimax M2.5/M2.7
    Calibration target for AI inference margin models
    Strengths
    Company feedback indicates M2.7 text/code inference gross margin can reach about 40%, supported by relatively high TPS assumptions.
    Weaknesses
    Achieving a 40% margin requires a high TPS ratio, and the model is sensitive to discount and workload changes.
    Comparison
    M2.5's single-GPU Lambda benchmark throughput is about 5.1x that of GLM-5, yet Z.ai model margins may still be higher.
    Risks
    Expanded commercial discounts and a higher share of text/code workloads causing the revenue mix to be less favorable than multimodal.
  • DeepSeek V4
    Observation target for China's AI frontier capability and price competition
    Strengths
    Close to the global SOTA of several months ago while costing less; open source/open weights help lift the overall capability of Chinese AI labs.
    Weaknesses
    The capability gap versus domestic peers is narrowing, and low-price strategy may weaken pricing power.
    Comparison
    V4-Pro pricing is close to domestic leading models such as GLM-5.1 and Qwen3.6 Max; V4-Flash is near the low end of lightweight low-cost model pricing.
    Risks
    Low prices and limited-time discounts intensify industry competition, and long-term profitability remains to be proven.

Key data

  • Minimax M2.7 base-case inference margin42.5%Based on model assumptions of a 20% commercial discount, TPS of 9,000, GPU utilization of 90%, and GPU cost of $1.6 per hour.
  • Z.ai GLM-5.1 base-case inference margin50.3%The report believes the GLM series may still achieve higher margins than Minimax models under comparable discount assumptions.
  • Z.ai GLM-5 base-case inference margin33.9%Based on the report's updated token economics model.
  • Minimax M2.5 base-case inference margin29.5%Base case derived after the report raised its H1 2026 TPS assumption.
  • GPU cost assumption$1.6/GPU-hourCompany feedback suggests this estimate is broadly reasonable, and the report assumes leading Chinese AI labs have broadly similar GPU costs.
  • Base commercial discount assumption20%The report adds a commercial discount variable and notes that customer mix and limited-time promotions can materially affect revenue per token.
  • Lambda.ai throughput benchmarkM2.5 about 4,031 tok/s/GPU; GLM-5 about 788 tok/s/GPUAccording to the report, M2.5 single-GPU throughput is about 5.1x that of GLM-5.
  • Tencent target priceHK$780Based on 20x FY+1 PE valuation; rated Outperform.
  • Alibaba target priceUS$180/HK$176Based on SOTP valuation of FY+1 revenue and profit for the core e-commerce and cloud businesses; rated Outperform.

Impact & implications

The investment implication is that Chinese AI model capabilities and commercialization competition are still evolving rapidly, and inference margins cannot be judged by list prices alone, but must also take into account TPS, discounts, workloads, and GPU utilization. For internet leaders, frontier model capabilities, the speed of AI infrastructure iteration, application-layer distribution entry points, and commercialization efficiency will shape the medium- to long-term valuation narrative; however, price wars in low-cost models, commercial discounts, and regulatory/macro risks will weigh on earnings visibility.

Risks

  • Intensifying price competition among low-cost lightweight AI models, reducing revenue per million tokens.
  • Commercial discounts, customer mix, and limited-time promotions may cause external margin estimates to diverge from actual financial results.
  • Key assumptions such as TPS, GPU utilization, and KV cache hit rates are difficult to verify precisely from the outside.
  • China's internet leaders still face risks from macro consumption, regulation, platform competition, and fluctuations in user engagement.
  • How AI lab costs are classified among cost of sales, R&D expense, and marketing expense may affect comparability.

What to watch

  • Actual inference margins and discount levels disclosed later by AI labs such as Minimax and Z.ai.
  • Follow-up adoption of DeepSeek V4, pricing strategy, and potential financing arrangements.
  • The release cadence of subsequent official versions and larger models after Tencent hy3-preview.
  • Whether Tencent's reorganization of AI data and infrastructure leads to faster model iteration.
  • Changes in the workload mix of multimodal, video generation, text-to-speech, and text/code tasks.
  • Commercialization progress of AI among Chinese internet companies in advertising, gaming, the WeChat ecosystem, and cloud businesses.
Zhejiang ICP No. 2022035445-5
Disclaimer: Market data, charts, indicators, research views, and other information provided on this website are intended solely for information display, research communication, and educational reference. They should not be regarded as personalized investment advice, securities recommendations, trading instructions, solicitations, or guarantees of return. While we strive to improve the reliability of our data and content, such information may still be subject to delays, errors, incompleteness, or untimely updates due to source differences, methodological limitations, system processing, or market volatility. Users should exercise independent judgment based on their own circumstances and bear all risks and responsibilities arising from the use of this website.

Settings

Sign in to view recent logins