Quick Summary
Covering the latest research from top Wall Street investment banks

Kimi K3 exceeded expectations in capability, and the cost-efficiency advantage of Chinese AI models continues to strengthen

Institution
Jefferies
Date
2026-07-16
Authors
Thomas Chong, Zoey Zong
Company
-
Ticker
-
Industry
Large AI Models and Internet
Rating
-
BullishLow confidenceThe report believes that Kimi K3's model capabilities and cost efficiency exceeded market expectations, that the intelligence gap between Chinese and U.S. models is narrowing, and that rising token consumption benefits cloud service providers, AI labs, and the application ecosystem.
AuthorsThomas Chong, Zoey Zong
CoverageOther
Business segmentsLarge Language Models、AI Agent、Cloud Services、Video Generation Models、Model API Pricing
Research firm divisions/subsidiariesJefferies(Other)

AI summary card

Kimi K3 exceeded expectations in capability, and the cost-efficiency advantage of Chinese AI models continues to strengthen

Jefferies believes that the open 2.8T model Kimi K3 launched by Moonshot AI ranks near the top across multiple benchmarks and reinforces the investment thesis around Chinese models in intelligence catch-up, API costs, and token consumption growth.

This report is an industry thematic tracker and does not provide a rating, target price, or current share price for any single company. The overall stance is positive, with a focus on bullishness toward cost-effective Chinese models and the cloud services and AI application chain.
Kimi K3Chinese AI ModelsModel API PricingOpenRouter Token ConsumptionCloud Service Providers BenefitingAI Agent
  • Kimi K3 is described as the first open 2.8T model, ranking No. 3 in the Artificial Analysis Index and No. 1 in the Frontend Code Arena.
  • Kimi K3 is designed for long-cycle coding, knowledge work, and reasoning, with 1M context and native multimodal capabilities.
  • The report emphasizes that Chinese models such as DeepSeek, Qwen, Kimi, MiniMax, and Zhipu have strong cost-efficiency advantages over U.S. models, with API costs only a fraction of those of U.S. models.
  • OpenRouter data shows that total token consumption in the week of July 6 to July 12 rose 12.6% week over week to 52.6tn, with Chinese model token consumption at 27.6tn, higher than the 6.3tn of U.S. models.
  • Beneficiary directions include Baidu, Alibaba, Tencent, Kingsoft Cloud, AI labs, and Kling, with the core logic being that rising demand for AI models and Agents drives computing power, cloud, and the application ecosystem.

Report interpretation

Overview

This report is the 58th installment of Jefferies' AI Series, tracking model capabilities, cost performance, global benchmark rankings, token consumption, and model pricing following Moonshot AI's release of Kimi K3. The report views the release of Kimi K3 as an unexpected positive surprise for the market, reflecting the rapid iteration of Chinese models in large parameter scale, long context, coding ability, Agent capability, and cost efficiency.

Core views

The core view is that the capability gap between Chinese large models and leading U.S. frontier models is narrowing, while Chinese models are showing stronger relative advantages in API cost, architectural efficiency, and token consumption growth. Kimi K3 ranks near the top in Artificial Analysis, Frontend Code Arena, and dimensions such as coding, general Agents, and vision Agents; Chinese models such as DeepSeek, Qwen, Kimi, MiniMax, and Zhipu gain cost advantages through MoE, attention mechanisms, and higher model compute utilization. On the demand side, OpenRouter token consumption is rebounding, with Chinese model consumption in particular exceeding that of U.S. models, creating a positive catalyst for ecosystem participants such as BAT, Kingsoft Cloud, AI labs, and Kling.

Analysis framework

The report uses an industry data-tracking approach, combining model benchmark tests, API input/output pricing, OpenRouter token consumption, the Token Expenditure Index, model architecture descriptions, and cross-country comparisons between Chinese and U.S. models to assess the competitiveness of Kimi K3 and the Chinese AI model ecosystem.

Methodology notes

  • Model Capability BenchmarkArtificial Analysis Intelligence Index v4.1

    Multi-dimensional intelligence evaluation

    This index integrates 9 evaluations including GDPval-AA v2, t-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR to compare the overall intelligence level of leading global frontier models.

  • Usage Demand TrackingOpenRouter token consumption

    Token consumption volume

    The report uses OpenRouter's weekly token consumption to measure model usage intensity and compares the usage rankings of Chinese models, U.S. models, and specific models.

  • Cost Efficiency TrackingToken Expenditure Index

    Expenditure index weighted by model usage mix

    The index rises as user demand shifts toward higher-priced models, and the report uses it to observe whether model demand is moving toward premium or higher-cost services.

  • Price ComparisonModel Input/Output API Pricing

    Input and output price per million tokens

    The report compares the API input and output prices of Kimi K3 with global models to assess cost performance at comparable capability.

Asset mapping & comparison

Structured mapping from thesis to named assets (strengths, weaknesses, peers, risks).

  • Moonshot AI / Kimi K3
    Core research subject
    Strengths
    It features an open 2.8T parameter scale, 1M context, native multimodality, strong coding and Agent capabilities, and ranks near the top in multiple benchmark tests.
    Weaknesses
    The report does not provide data on commercialization revenue, user retention, or profitability, and the full release of model weights still needs to wait until July 27.
    Comparison
    It ranks No. 3 in the Artificial Analysis Index, behind Fable5 and GPT-5.6 Sol(Max), but ranks No. 1 in the Frontend Code Arena.
    Risks
    If subsequent real developer adoption, API stability, or the effect of open model weights falls short of expectations, the market's pricing of its capability breakthrough may pull back.
  • DeepSeek、Qwen、Kimi、MiniMax、Zhipu
    Chinese high-cost-performance model camp
    Strengths
    Through architectural optimizations such as MoE, GQA, sparse attention, linear attention, and MLU, their API costs are lower than those of U.S. models, while the intelligence gap continues to narrow.
    Weaknesses
    Top global model capabilities are still importantly led by U.S. companies such as Anthropic, Google, and OpenAI, and some high-end benchmarks still leave room for catch-up.
    Comparison
    The report says Chinese model API costs are only a fraction of those of U.S. models, and cites the Stanford AI Index showing that top U.S. models lead Chinese models by 2.7%.
    Risks
    Price competition may compress profit margins for model services, while regulation, compute supply, and access to overseas ecosystems remain uncertain factors.
  • Baidu、Alibaba、Tencent、Kingsoft Cloud
    Potential beneficiary cloud service providers and internet platforms
    Strengths
    Growth in the use of AI models and Agents will increase demand for computing power, cloud services, and model hosting, and major platforms have advantages in infrastructure and customer base.
    Weaknesses
    The report does not provide revenue sensitivity, profit elasticity, or rating changes for individual companies.
    Comparison
    Compared with pure model companies, cloud service providers may benefit more directly from growth in token consumption and enterprise workflow deployment.
    Risks
    If token consumption growth does not translate into paid API or cloud revenue, or if enterprises set token consumption caps for employees, earnings realization may fall short of expectations.
  • Kling and the video generation model ecosystem
    Beneficiary direction from AI application demand
    Strengths
    Video generation model pricing and membership plans are included in the report's tracking, showing that multimodality and content generation remain important application-layer scenarios.
    Weaknesses
    The pricing table information for video generation in the report is relatively fragmented and lacks clear market share and profitability metrics.
    Comparison
    Compared with general LLMs, video generation applications rely more on per-video cost, membership conversion, and content creation demand.
    Risks
    Price competition in video generation, promotional discounts, and falling model costs may affect commercialization quality.

Key data

  • Kimi K3 parameter scale2.8TThe report states that Moonshot AI released Kimi K3, the first open 2.8T model.
  • Kimi K3 Artificial Analysis rankingNo. 3Ranked behind Fable5 and GPT-5.6 Sol(Max).
  • Kimi K3 Frontend Code Arena rankingNo. 1The report believes this ranking highlights its frontend coding capability.
  • Kimi K3 context1M contextIt also has native multimodal capabilities.
  • Delta Attention decoding efficiencyup to 6.3xThe report says it can achieve up to 6.3x faster decoding in a million-token context.
  • Attention Residuals training efficiency25% higherAdditional cost is below 2%.
  • Overall scaling efficiency2.5x improvement versus K2The report attributes this to architectural updates and the Stable LatentMoE framework.
  • OpenRouter total token consumption52.6tnWeek over week growth of 12.6% in the week of July 6 to July 12.
  • Chinese model token consumption27.6tnIn the week of July 6 to July 12, up 17.7% from the previous week.
  • U.S. model token consumption6.3tnThe report points out that Chinese models exceeded U.S. models in OpenRouter token consumption.
  • Token Expenditure Index1.57-1.61Range for the week of July 12, below 1.63-1.65 for the week of July 5 and also below 2.04 on May 31.
  • OpenRouter model token rankingHy3(free) 6.13T、MiMo-V2.5 5.95T、DeepSeek V4 Flash 5.22TThese were the top three in the latest weekly ranking.
  • Stanford AI Index lead gap2.7%The report cites the 2026 Stanford AI Index Report as saying that top U.S. models lead Chinese models by 2.7%.

Impact & implications

The investment implication is that the release of Kimi K3 strengthens the visibility of Chinese AI models in competition among leading global frontier models, and its cost-efficiency advantage may drive more developers, enterprises, and Agent workflows to adopt Chinese models. Rising token consumption will boost demand for cloud computing resources, model API calls, AI Agents, coding workflows, and content generation, benefiting cloud service providers, model labs, and application-layer companies such as video generation.

Risks

  • Benchmark performance of models may not fully represent stability, latency, developer experience, and enterprise deployment effectiveness in real production environments.
  • The cost advantage of Chinese models may trigger more intense price competition, causing API revenue growth and margin improvement to move out of sync.
  • Enterprise caps on employee token consumption may suppress spending on premium models and the rebound of the Token Expenditure Index.
  • Rapid iteration in model capabilities may shorten the leadership window of any single model, and Kimi K3's market surprise requires follow-up validation through usage volume and ecosystem adoption.
  • Comparisons between Chinese and U.S. models involve differences in data sources, evaluation frameworks, and samples, and ranking changes may lead to volatility in market expectations.

What to watch

  • Developer adoption and real application feedback after the full Kimi K3 model weights are released on July 27.
  • Whether OpenRouter's weekly token consumption continues to grow, and how the share of Chinese models changes relative to U.S. models.
  • Whether the Token Expenditure Index continues to weaken, or rebounds as demand for premium models recovers.
  • Ranking changes of Chinese models such as DeepSeek, Qwen, Kimi, MiniMax, and Zhipu in Artificial Analysis and in coding, Agent, and vision tasks.
  • Whether companies such as Baidu, Alibaba, Tencent, and Kingsoft Cloud reflect cloud revenue contributions from rising AI token consumption in earnings reports or operating data.
  • Price adjustments for model APIs and video generation services, real paid demand after promotions end, and unit economics.
Zhejiang ICP No. 2022035445-5
Disclaimer: Market data, charts, indicators, research views, and other information provided on this website are intended solely for information display, research communication, and educational reference. They should not be regarded as personalized investment advice, securities recommendations, trading instructions, solicitations, or guarantees of return. While we strive to improve the reliability of our data and content, such information may still be subject to delays, errors, incompleteness, or untimely updates due to source differences, methodological limitations, system processing, or market volatility. Users should exercise independent judgment based on their own circumstances and bear all risks and responsibilities arising from the use of this website.

Settings

Sign in to view recent logins