Report Interpretation
Covering the latest research from top Wall Street investment banks
Report InterpretationHilo Research

Artificial intelligence Report Interpretation

Citi argues that early recursive-self-improvement signals, falling inference costs and better application harnesses are broadening the AI intelligence flywheel. It maintains that safety-focused compute allocation is unlikely to alter its global AI-related CapEx projections of $1T in 2026 and $1.6T in 2027.

InstitutionCitigroup
Date20260918
IndustryArtificial intelligence

Summary

Citi argues that early recursive-self-improvement signals, falling inference costs and better application harnesses are broadening the AI intelligence flywheel. It maintains that safety-focused compute allocation is unlikely to alter its global AI-related CapEx projections of $1T in 2026 and $1.6T in 2027.

No subject-specific rating or target price.
Artificial intelligenceInferenceRecursive self-improvementCompute demandAI infrastructureModel releasesAI safetyAgent harnesses
  • Average realized cost-to-serve was down 54% from its May 28 peak and 43% since July 1.
  • Z.ai's Infra Agent moved GLM-5.3-Flash into production in under two weeks and tripled throughput across more than 100,000 accelerators.
  • A provider adapter harness lifted ARC-AGI-3 performance to 99.95% from 62.71% while cutting cost by 28%.
  • Only one proprietary and three open models launched during the week, but trailing-90-day releases remained at 68 proprietary and 40 open models.

Report Interpretation

Overview

This Citi AI-industry update examines whether recursive self-improvement, declining inference costs and application-layer advances can compound realized AI intelligence and demand for infrastructure. The report's central conclusion is that safety and alignment work may redirect compute but is more likely to add to total compute requirements than reduce the AI investment cycle.

Core views

Citi identifies early evidence of recursive self-improvement (RSI) as a potential new source of compute demand. Google’s Dream-RSI and related developments suggest that models may increasingly contribute to developing their successors. Citi notes that the debate is shifting from the timing of model releases toward the allocation of scarce compute. Meta’s reported delay of Muse to prioritize safety, and comments that much of its compute serves users rather than an RSI race, illustrate that frontier development is being balanced against safety, auditability and alignment. Citi nevertheless argues that such reallocation should not change its estimates for global AI-demand-related CapEx of $1T in 2026 and $1.6T in 2027. The report views infrastructure automation as an important bridge toward RSI. Z.ai’s Infra Agent reportedly put GLM-5.3-Flash into production in under two weeks, tripled throughput across more than 100,000 accelerators, and helped process more than 62T tokens in six days. While Citi stresses that the tool remains far from RSI, it believes autonomous optimization can speed model development while raising needs for inference and governance infrastructure. The delayed release of GLM-5.3 weights because of advanced cybersecurity capabilities, together with discussion at the AI Infra Summit, supports Citi’s view that allocating compute to safety and alignment could increase overall compute requirements. Lower deployment costs and improved efficiency are presented as complementary demand drivers. An Unsloth quantization of GLM-5.3 retained about 81% accuracy while making the model 83% smaller, reducing on-premise memory and accelerator needs. Citi argues that lower deployment barriers can widen on-premise AI adoption, improve control over sensitive data and help establish the conditions for RSI at scale. This is consistent with Ramp’s reported month-on-month adoption gains in six of seven sectors. Separately, average realized cost-to-serve fell 54% from its May 28 peak, 43% since July 1 and 2% week on week, which Citi views as a tailwind to both internal and external usage growth. Model-release cadence slowed in the latest week: one proprietary and three open models launched, versus five and six respectively in the prior week, and three and six four weeks earlier. Yet the trailing 90-day count reached 68 proprietary and 40 open releases; proprietary releases were up 31% month on month after an early-September surge, while open releases were flat. Citi also reports that roughly 69% of releases since August 1 set a new intelligence high for the relevant provider. It flags unconfirmed speculation of a frontier-model release near month-end as a possible test of whether a capability gain can redirect premium usage despite wider use of model routing. Finally, Citi argues that harness design is becoming a major determinant of realized intelligence rather than merely the underlying model. On the interactive ARC-AGI-3 benchmark, adding a provider adapter harness to the same model increased performance by 37 points to 99.95% from 62.71% and reduced cost by 28% to $18.8K from $26.1K. Salesforce similarly raised average success across seven enterprise-agent tasks to 78.0% from 29.2% by evolving its agent harness without changing the underlying model. Citi concludes that an increasing share of practical intelligence gains may arise from harnesses, strengthening the application layer as a source of differentiation for providers and enterprises.

Analysis framework

Citi combines weekly tracking of model releases, intelligence benchmarks, pricing and usage indicators with case studies of AI infrastructure, deployment optimization and agent-harness design. It links these signals to compute allocation, adoption and AI-related capital-expenditure implications.

Methodology notes

  • Industry AnalysisSupply-demand framework

    AI compute supply-demand and cost-to-serve analysis

    Citi uses changes in serving costs, compute allocation, model releases and infrastructure efficiency to assess how AI demand and compute requirements may evolve.

  • Industry AnalysisUpstream-Midstream-Downstream Transmission

    AI infrastructure-to-deployment transmission

    The report connects infrastructure automation and lower model memory requirements with wider deployment, inference usage, governance needs and demand for AI compute.

Key data

  • Global AI-demand-related CapEx$1T in 2026; $1.6T in 2027Citi's projection, which it says should remain intact despite greater compute allocation to safety and alignment.
  • Average realized cost-to-serveDown 54% from the May 28 peak, 43% since July 1, and 2% WoWCiti identifies lower serving costs as a usage-growth tailwind.
  • Z.ai Infra Agent throughputTripled across 100K+ accelerators; 62T+ tokens processed in 6 daysThe system moved GLM-5.3-Flash to production in under two weeks.
  • GLM-5.3 quantization~81% accuracy; 83% smallerCiti cites the Unsloth quantization as reducing on-premise memory and accelerator requirements.
  • Latest model launches1 proprietary and 3 openDown from 5 and 6 week on week; trailing-90-day totals were 68 proprietary and 40 open.
  • ARC-AGI-3 harness result99.95% from 62.71%; $18.8K from $26.1KA provider adapter harness added 37 performance points and reduced cost by 28% using the same model.
  • Salesforce enterprise-agent task success78.0% from 29.2%Improvement across seven tasks came from evolving the agent harness rather than changing the underlying model.

Impact & implications

Citi’s analysis suggests that efficiency gains need not reduce AI infrastructure demand: lower costs can broaden usage and deployment, while RSI-oriented development and safety, alignment and governance work can add compute requirements. It also highlights harness design as an increasingly important source of competitive differentiation at the application layer.

Risks

  • Advanced cybersecurity capabilities may delay model-weight releases, as cited in the GLM-5.3 example.
  • Safety, auditability and alignment requirements may change the pacing and allocation of frontier-model compute.

What to watch

  • Whether a frontier-provider model release near month-end produces a capability gain sufficient to redirect premium usage despite model routing.
  • How providers allocate compute among user serving, model development, safety, auditability and alignment.
  • Further changes in inference serving costs, model-release cadence and deployment efficiency.
  • Upcoming AI and infrastructure events, including Meta Connect, OpenAI DevDay, OCP Global Summit and NVIDIA GTC.
Zhejiang ICP No. 2022035445-5
Disclaimer: Market data, charts, indicators, research views, and other information provided on this website are intended solely for information display, research communication, and educational reference. They should not be regarded as personalized investment advice, securities recommendations, trading instructions, solicitations, or guarantees of return. While we strive to improve the reliability of our data and content, such information may still be subject to delays, errors, incompleteness, or untimely updates due to source differences, methodological limitations, system processing, or market volatility. Users should exercise independent judgment based on their own circumstances and bear all risks and responsibilities arising from the use of this website.

Settings

Sign in to view recent logins