Quick Summary
Covering the latest research from top Wall Street investment banks

AI inference profits are not just about cost; the revenue side and workload mix are the key differentiators

Institution
Bernstein
Date
2026-04-21
Authors
Robin Zhu, Charles Gou, Min-Joo Kang
Company
-
Ticker
-
Industry
Internet; Artificial Intelligence
Rating
Tencent: Outperform; Alibaba: Outperform
NeutralLow confidenceThe report argues AI inference unit economics will increasingly matter for competitiveness. It expects Z.ai margins to improve after GLM-5/GLM-5.1 price increases and sees Qwen/Alibaba Cloud retaining a compute-cost edge, while Minimax faces potential Rev/Mtok pressure as workload mix shifts away from high-priced text-to-speech.
AuthorsRobin Zhu, Charles Gou, Min-Joo Kang
Target priceTencent HK$780; Alibaba ADR US$180; Alibaba 9988.HK HK$176
Asset classesEquity
SubsidiariesAlibaba Cloud、Qwen
Business segmentsAI inference、Cloud computing、AI models、Internet platforms、Text/coding workloads、Text-to-speech、Video generation
Research firm divisions/subsidiariesBernstein(Other)

AI summary card

AI inference profits are not just about cost; the revenue side and workload mix are the key differentiators

Bernstein believes that Chinese AI labs' inference margins are jointly determined by revenue per million tokens, KV cache hit rates, input-output ratios, modality mix, and GPU hourly costs; Z.ai is expected to improve margins in H1 2026, Qwen has a cost advantage backed by Alibaba Cloud, while Minimax may be dragged down as workloads shift from high-priced speech to text/code and agent orchestration.

The coverage table shows both Tencent and Alibaba rated Outperform; target prices are HK$780 for Tencent, US$180 for Alibaba ADR, and HK$176 for Alibaba 9988.HK, respectively.
China InternetAI inference economicsToken pricingRev/MtokGPU costQwenZ.aiMinimax
  • Revenue-side variables may have a greater impact on real-world AI lab inference margins than cost-side variables, especially differences in Rev/Mtok across modalities and workloads.
  • Text-to-speech token pricing is significantly higher than text/code and video generation, so Minimax's previously high Open Platform gross margin may have partly benefited from audio APIs.
  • Price increases for Z.ai's GLM-5 and GLM-5.1 are expected to offset cloud vendors' compute price hikes and drive inference margin improvement in H1 2026.
  • Backed by Alibaba Cloud compute and the profit chain of a hyperscale cloud provider, Qwen may enjoy a more durable unit compute-cost advantage than independent AI labs.

Report interpretation

Overview

This report is a deep-dive study on China Internet and AI inference economics, seeking to explain the formation mechanism of AI lab inference margins with a relatively low technical barrier. It defines direct inference gross profit as revenue per token minus cost per token, excluding company-level expenses and training costs, and relies mainly on Minimax and Z.ai financial disclosures, public token pricing, OpenRouter data, GPU hourly cost estimates, and cloud vendor channel information. The core question is not whether AI revenue is growing, but how different models, modalities, regional pricing, cache hit rates, input-output lengths, and GPU utilization jointly determine unit economics.

Core views

The report's core views include: first, different AI inference tokens are not equivalent, and the revenue intensity of text/code, video generation, and text-to-speech varies greatly; second, differences in revenue per million tokens may be more decisive for real-world margins than cost-per-token optimization; third, due to price increases for Z.ai, GLM-5, and GLM-5.1, Z.ai's inference margins are likely to improve meaningfully in H1 2026; fourth, although Minimax has disclosed high Open Platform gross margins, its blended Rev/Mtok may come under pressure if workloads shift from high-priced text-to-speech toward lower-priced text/code and agent orchestration; fifth, Qwen/Alibaba Cloud has a medium-term cost advantage over independent AI labs due to its cloud infrastructure and compute supply chain.

Analysis framework

The report uses a combined top-down and bottom-up 'rough but explainable' model: on the revenue side, it estimates Rev/Mtok after adjusting for input tokens, output tokens, KV cache hit rate, cached token discounts, domestic vs. international pricing differences, and multimodal price conversions; on the cost side, it estimates cost per GPU hour, tokens-per-second throughput, GPU utilization, model parameter scale, batching efficiency, and cloud vendor markups. It then uses sensitivity analysis across different models, regions, modalities, and workloads for Z.ai and Minimax to assess directional differences in inference margins.

Methodology notes

  • unit_economicsRev/Mtok model

    revenue per million tokens

    The revenue side is jointly determined by the input-output token ratio, input/output token pricing, KV cache hit rate, and cached-token discounts; for multimodal models, video seconds, audio character counts, and similar units must also be converted into equivalent token revenue.

  • unit_economicsCost/Mtok model

    cost per million tokens

    The cost side is derived from GPU hourly cost, per-GPU tokens-per-second throughput, GPU utilization, and seconds per hour; model architecture, MoE sparsity, batching efficiency, and cloud vendor service markups all affect the result.

  • scenario_analysisworkload mix sensitivity

    workload mix sensitivity

    The report compares different workloads such as text/code, video, text-to-speech, and agent orchestration, emphasizing that changes in modality mix can alter blended Rev/Mtok and thus affect AI lab gross margins.

  • market_comparisonvaluation comparison

    internet company valuation comparison

    The report includes valuation comparison tables for Chinese, Asian, and US internet companies, providing coverage-company and peer context through P/E, EV/sales, market cap, ratings, target prices, and latest prices.

Asset mapping & comparison

Structured mapping from thesis to named assets (strengths, weaknesses, peers, risks).

  • Alibaba / Qwen / Alibaba Cloud
    benefits from cloud compute cost advantages and Qwen's ability to capture multimodal demand
    Strengths
    Cloud infrastructure scale, compute scheduling, the ability to raise prices to external AI customers, and the Qwen model ecosystem may lead to lower Cost/Mtok.
    Weaknesses
    Compute constraints, chip import restrictions, and cloud business capex pressure may still affect profit realization.
    Comparison
    Relative to independent AI labs, Alibaba Cloud as a hyperscaler is more likely to retain an advantage in GPU hourly costs and the service chain.
    Risks
    Macro consumption, platform competition, regulation, losses in innovation businesses, and AI price competition may weaken valuation and earnings elasticity.
  • Z.ai
    beneficiary of improving inference margins
    Strengths
    Price increases for GLM-5 and GLM-5.1, general reasoning workloads, and international pricing premiums help lift Rev/Mtok.
    Weaknesses
    Limited disclosure, cost-side dependence on external GPU hourly cost assumptions, and domestic pricing below international pricing.
    Comparison
    The report estimates Z.ai's text/code inference margins are higher than Minimax's unless Minimax's throughput advantage is far greater than assumed.
    Risks
    Cloud vendor price hikes, errors in model throughput assumptions, and changes in KV cache and ISL:OSL assumptions could alter the conclusion.
  • Minimax
    its high-margin sustainability faces pressure
    Strengths
    M2.5/2.7 is positioned as a lightweight agent backbone; the model is smaller and sparser, which may support higher throughput and GPU utilization; Open Platform gross margin was high in 9M 2025.
    Weaknesses
    No price increases, lower text/code Rev/Mtok; if the share of high-priced text-to-speech workloads declines, overall revenue intensity will come under pressure.
    Comparison
    Compared with Z.ai, Minimax may be better on throughput and utilization, but the report believes this is insufficient to fully offset the revenue-side gap.
    Risks
    Migration of workloads toward OpenClaw, text/code, and agent orchestration may compress blended Rev/Mtok and gross margins.
  • Tencent
    coverage company, related to internet valuation and the AI theme
    Strengths
    Rated Outperform in the table, with a target price of HK$780, implying upside from the current price of HK$522.50.
    Weaknesses
    The report is not primarily a dedicated analysis of Tencent's AI inference margins, serving more as China Internet coverage and valuation context.
    Comparison
    Within China Internet coverage, it is shown alongside Alibaba, PDD, Meituan, NetEase, JD, and others in valuation comparisons.
    Risks
    Macro credit and retail consumption, fluctuations in user engagement, competition in gaming and advertising, and regulatory risks such as antitrust.

Key data

  • Z.ai API gross margin22.4% in H2 2025, versus -11.7% a year earlier, implying a year-over-year incremental margin of 30.3%The report believes GLM-5/GLM-5.1 price increases could drive further improvement in H1 2026.
  • Minimax Open Platform gross margin69.4% in 9M 2025, versus 62.3% in 9M 2024, implying a year-over-year incremental margin of 73.9%The company did not disclose a full-year 2025 figure; the report worries that the high margin benefited from high text-to-speech pricing and may be affected by future workload migration.
  • Text-to-speech Rev/MtokThe report estimates Minimax's text-to-speech list price may exceed US$10/Mtok, and around US$17/Mtok on an HD basisThis is significantly higher than text/code tokens, explaining the large impact of multimodal workloads on margins.
  • Z.ai text/code Rev/MtokAfter adjusting for typical input-output ratios and KV cache, blended Rev/Mtok is roughly in the US$1 rangePrice increases for GLM-5 and GLM-5.1 are an important driver of margin improvement in H1 2026.
  • Minimax text/code Rev/MtokText/code tokens for the lightweight model are roughly in the US$0.2/Mtok rangeIf the share of agent workloads such as OpenClaw increases, overall revenue intensity may decline.
  • KV cache hit rateCharts show a median of about 39% for GLM-5 and about 71% for Minimax-M2.5Cache hit rates affect effective input-token revenue and also reflect differing workload characteristics.
  • GPU hourly costBased on Alibaba Cloud AI revenue, compute scale, and H20 server leasing quotes from social media, the report estimates that a range below US$1/hour is reasonableAlibaba Cloud is estimated at about RMB7-9/hour, while bare-metal server leasing is about RMB7.0-8.5/hour, with the latter looking more like a lower bound.
  • Alibaba Cloud external AI revenue assumptionAbout RMB9.0bn last quarter, or about US$1.3bnThe report uses this to back out an assumption of about 1.2-1.3GW of compute capacity and roughly 2.5 million GPUs.
  • Coverage company ratings and target pricesTencent O, target price HK$780; Alibaba O, ADR target price US$180, 9988.HK target price HK$176Current prices in the table are Tencent HK$522.50, BABA US$141.01, and 9988.HK HK$137.00.

Impact & implications

The investment implication is that AI commercialization should not be judged only by revenue growth or model popularity, but by whether inference unit economics are sustainable. For internet platforms, companies with cloud infrastructure and GPU procurement and scheduling capabilities are more likely to build moats on the cost side; for independent AI labs, short-term high margins may come from favorable modality mix or pricing windows, but as the share of agent and text/code workloads rises, revenue intensity may be diluted. The report therefore favors the relative cost position of Qwen/Alibaba Cloud, sees high visibility on near-term margin improvement for Z.ai, and remains cautious on the sustainability of Minimax's high margins.

Risks

  • AI lab disclosure is limited, and many cost and throughput assumptions can only be estimated through public prices, channel information, and rough models.
  • GPU hourly cost, utilization, tokens-per-second throughput, batching efficiency, and model architecture differences may significantly change Cost/Mtok estimates.
  • Commercial discounts are unobservable, so public OpenRouter or list prices may overstate true Rev/Mtok.
  • The competitive landscape and customer willingness to pay for high-priced text-to-speech may be unsustainable.
  • If workload mix quickly shifts toward lower-priced text/code or agent orchestration, blended Rev/Mtok for platforms such as Minimax may decline.
  • Cloud vendor compute price hikes, chip import restrictions, GPU supply, and domestic-vs.-international pricing differences may affect AI lab margins.
  • Tencent and Alibaba still face traditional internet risks such as macro conditions, platform competition, regulation, and losses in innovation businesses.

What to watch

  • Whether Z.ai's next earnings disclosure shows a meaningful improvement in API gross margin and incremental margin.
  • Whether domestic and international pricing for GLM-5, GLM-5.1, and subsequent models continues to rise, and the extent of commercial discounts.
  • Whether Minimax begins to raise prices, or continues to use M2.5/2.7 as a low-cost agent backbone.
  • Changes in the share of agent workloads such as OpenClaw and Hermes agents, and their impact on ISL:OSL, KV cache, and Rev/Mtok.
  • Whether text-to-speech API pricing maintains a high premium or falls back as competition intensifies.
  • Alibaba Cloud's external AI revenue, compute utilization, GPU hourly cost, and Qwen's closed-source/commercialization strategy.
  • Available GPU supply, rental pricing, and China AI compute restrictions for H20/H100/H800 and similar chips.
Zhejiang ICP No. 2022035445-5
Disclaimer: Market data, charts, indicators, research views, and other information provided on this website are intended solely for information display, research communication, and educational reference. They should not be regarded as personalized investment advice, securities recommendations, trading instructions, solicitations, or guarantees of return. While we strive to improve the reliability of our data and content, such information may still be subject to delays, errors, incompleteness, or untimely updates due to source differences, methodological limitations, system processing, or market volatility. Users should exercise independent judgment based on their own circumstances and bear all risks and responsibilities arising from the use of this website.

Settings

Sign in to view recent logins