ZCube's new network architecture may improve Z.ai inference efficiency
AI summary card
ZCube's new network architecture may improve Z.ai inference efficiency
Morgan Stanley believes that by optimizing KV cache traffic between the inference prefill and decode stages, ZCube can reduce network-related capital expenditures and improve GPU inference throughput, which is important for Z.ai's ARR upside potential.
- ZCube can reduce switch and optical module capital expenditures by 33% and increase GPU inference throughput by 15%.
- In the GLM-5.1 coding test, ZCube reduced TTFT P99 latency by 40.6%.
- The report believes that China's AI path places greater emphasis on scaling models and inference through algorithmic and engineering efficiency innovation.
- The valuation method uses DCF, assuming a 15% WACC and a 3% perpetual growth rate, and the target price implies a 2027 38x P/S multiple.
Report interpretation
Overview
This report focuses on the Z.ai technology progress related to KnowledgeAtlasTechnologyJSCLtd (2513.HK). Z.ai launched the new ZCube network architecture on 2026-05-21, and Morgan Stanley views it as an important innovation to improve large-model inference efficiency and alleviate bottlenecks in inference compute supply. The core view is that as model scale and inference scale expand in tandem, inference efficiency becomes increasingly important, and the network has become a key component of inference systems.
Core views
The report argues that the upside potential in Z.ai's ARR depends in large part on its inference computing capacity. ZCube uses a new network topology to optimize KV cache traffic between the inference prefill and decode stages, shifting inference efficiency from point optimization to system-level optimization. Specific effects include a 33% reduction in switch and optical module capital expenditures, a 15% increase in GPU inference throughput, and a 40.6% reduction in TTFT P99 latency in the GLM-5.1 coding test. This is consistent with the report's observation on China's AI development path, namely achieving scale through algorithmic innovation and improved engineering efficiency.
Analysis framework
This report is primarily a technical event commentary, linking the inference efficiency improvement brought by ZCube's launch to Z.ai's commercialization potential, ARR upside, and China's AI engineering-efficiency path; the valuation section uses DCF and supplements it with an implied 2027 P/S multiple as a cross-check.
Methodology notes
Discounted cash flow valuation
The report values the company using the DCF method, assuming a 15% WACC and a 3% perpetual growth rate.
Price-to-sales ratio
The report states that the target price implies a 2027 38x P/S multiple, used to assess the fit between the company's revenue scale and valuation.
System-level inference network optimization
The report evaluates ZCube's impact on inference efficiency from the angles of network topology, KV cache traffic, GPU throughput, TTFT P99 latency, and network equipment capex.
Asset mapping & comparison
Structured mapping from thesis to named assets (strengths, weaknesses, peers, risks).
- KnowledgeAtlasTechnologyJSCLtd(2513.HK)Report subject company / stock
- Strengths
- Z.ai launched the new ZCube network architecture, which the report believes can improve inference efficiency and ease constraints on inference compute supply.
- Weaknesses
- The report does not provide clear current profit, revenue, or ARR data, and the valuation is sensitive to future growth and technology execution.
- Comparison
- The report places ZCube within the path of China AI scaling through engineering efficiency and algorithmic innovation, rather than relying solely on brute-force compute expansion.
- Risks
- Geopolitical risk, intensifying competition and price wars, and model performance lagging peers.
- Z.ai / ZCubeCore technology and business driver
- Strengths
- Optimizes KV cache traffic and network topology, reduces network equipment capex, improves GPU throughput, and lowers TTFT P99 latency.
- Weaknesses
- The metrics disclosed in the report come from a specific GLM-5.1 coding test; it remains to be seen whether they can be generalized to broader models and real inference scenarios.
- Comparison
- Represents inference efficiency shifting from point optimization to system-level network optimization.
- Risks
- If model performance or commercialization lags peers, the improvements in technical efficiency may not fully translate into revenue growth.
Key data
- ZCube发布时间2026-05-21Z.ai launched the new ZCube network architecture on this date.
- 交换机和光模块资本开支变化-33%The report says ZCube can reduce related capex.
- GPU推理吞吐变化+15%The report says ZCube can improve GPU inference throughput.
- TTFT P99延迟变化-40.6%Based on the GLM-5.1 coding test.
- WACC假设15%Discount rate assumption in the DCF valuation method.
- 永续增长率假设3%Terminal growth assumption in the DCF valuation method.
- 目标价隐含估值2027年38倍P/SThe report does not provide an explicit target price in the extracted text.
Impact & implications
If ZCube's efficiency gains can be sustained in production, Z.ai may be able to expand inference services with lower network-related capital expenditures and higher GPU utilization efficiency, thereby easing compute supply constraints and supporting ARR growth. From an investment perspective, this technological progress strengthens the company's narrative in AI infrastructure and large-model service efficiency, but valuation still depends on future model performance, commercialization revenue realization, and changes in the competitive landscape.
Risks
- Geopolitical risk.
- Intensifying competition and price wars.
- Model performance lagging peers.
- Valuation is sensitive to a 15% WACC, a 3% perpetual growth rate, and future revenue growth assumptions.
- Morgan Stanley discloses that it may have business relationships with the covered company; investors should treat this research as only one factor in investment decisions.
What to watch
- Whether throughput, latency, and cost improvements from ZCube can be consistently reproduced in real production inference workloads.
- Whether Z.ai ARR growth improves as inference efficiency rises.
- Changes in the performance of GLM series models relative to global SOTA and Chinese peers.
- The competitive landscape of AI infrastructure and the intensity of price wars.
- The impact of geopolitical and supply-chain constraints on inference compute availability.