Quick Summary
Covering the latest research from top Wall Street investment banks

AI Inference Reshapes Memory Demand: KV Cache Offloading and Agentic AI Become Core Drivers

Institution
TrendForce
Date
20260616
Company
xprecision, Compass, Nvidia, Intel, Arm Holdings, Ampere, SK hynix, Arm
Ticker
BYTES, COMP, NVIDIA, INTEL, ARM, AMPERE, SKHYNIX
Industry
Software - Application, AI, DRAM, SSD, Computer Hardware, Artificial Intelligence, Semiconductor, Storage
Rating
BullishHigh confidenceMedium-termThe report is bullish on the structural incremental demand for CPU RAM and SSD storage brought by KV cache offloading and Agentic AI in the AI inference era.
CoverageUnited States、Other
Research firm divisions/subsidiariesTrendForce(Subsidiary/Legal Entity)

AI summary card

AI Inference Reshapes Memory Demand: KV Cache Offloading and Agentic AI Become Core Drivers

Hardware demand in the AI inference era is shifting from training to inference; KV cache offloading technology spawns new demand for SSD PODs, while the explosion of Agentic AI pushes the CPU-to-GPU compute ratio towards 1:1, significantly driving CPU RAM demand.

AI InferenceKV CacheSSD PODAgentic AICPU RAMNvidiaStorage Architecture
  • AI inference imposes three new requirements on hardware: higher QPS, longer context windows, and more inference steps.
  • Token output per question surges over 5x, marking the industry's entry into the Test-time Scaling phase.
  • KV cache occupancy explodes with context and Batch Size; Nvidia launches Dynamo software and CMX platform to offload KV cache to CPU RAM and SSD POD.
  • Agentic AI requires significant CPU for orchestration and tool invocation; the CPU-to-GPU workload ratio will shift from 1:4 or 1:8 to 1:1.
  • Nvidia Vera CPU supports 1.5TB LPDDR5X, but limited by LPDRAM capacity, the next-generation module memory capacity is forced to halve.
  • 2026 will see a comprehensive refresh of Agentic AI-exclusive CPUs, with Intel, AMD, Arm, and Ampere all launching new products.

Report interpretation

Overview

This research report, published by TrendForce, deeply analyzes the structural transformation brought to memory and storage systems in the AI inference era. The report points out that as AI moves from training to inference, the core contradiction of hardware demand is shifting. The explosive growth of KV Cache has spawned new demand for storage hierarchy offloading, while the rise of Agentic AI has completely reshaped the compute ratio between CPU and GPU, thereby opening up a huge incremental market for CPU RAM and enterprise SSDs.

Core views

Hardware demand in the AI inference phase differs fundamentally from the training phase, mainly reflected in higher Queries Per Second (QPS), longer context windows, and more inference steps. Nvidia data shows that since the second half of 2024, the average token output per question has surged at a rate of over 5x per year, reaching 30,000 to 40,000 tokens, marking the industry's entry into the 'Test-time Scaling' thinking phase, which directly translates into huge consumption of memory and compute. In terms of memory occupancy, model weights belong to static allocation, while KV cache is dynamically allocated. As conversation length and Batch Size increase, the memory consumption of KV cache rises sharply. When GPU HBM capacity is insufficient, the system has to discard KV cache and recalculate, leading to increased latency and higher Total Cost of Ownership (TCO). To this end, Nvidia launched Dynamo software to offload infrequently accessed KV cache to CPU RAM and SSDs, which have lower bandwidth but larger capacity and better cost efficiency. Meanwhile, Nvidia released the CMX Context Memory Storage platform in early 2026, adding a Pod-level context layer (SSD POD, i.e., G3.5 layer) between local SSD and shared storage via BlueField-4 DPU, specifically designed to manage the massive KV cache generated by long-context workloads. On the other hand, the explosion of Agentic AI is reshaping CPU architecture. Agents need to actively plan, invoke tools, and execute operations, which requires the CPU to undertake tasks such as orchestration, data routing, and sub-agent evaluation. Therefore, the CPU-to-GPU workload ratio is evolving from the traditional 1:4 or 1:8 towards 1:1. Nvidia launched the Vera CPU designed specifically for agents, supporting up to 1.5TB of LPDDR5X memory, three times that of the previous generation Grace. However, due to insufficient supplier LPDRAM capacity allocation in 2027, Nvidia was forced to halve the SOCAMM memory capacity of the next-generation Vera Rubin superchip module. Besides Nvidia, Intel, AMD, Arm, and Ampere are also launching new products intensively in 2026, welcoming the comprehensive CPU replacement wave brought by Agentic AI.

Analysis framework

The report adopts an analytical framework of 'Demand Breakdown and Architecture Evolution'. First, the institution breaks down the hardware demand of AI inference into two dimensions: model weights (static) and KV cache (dynamic). By quantifying the surge trend of Token output, it deduces the HBM capacity bottleneck, thereby引出 (leading to) the technical path of 'KV cache offloading' and its driving effect on SSD POD and CPU RAM. Second, starting from the workflow characteristics of Agentic AI (planning, tool invocation), the institution deduces its dependence on low latency and CPU compute, thus concluding the industry trend of CPU/GPU ratio converging to 1:1, and finally landing on the incremental demand for CPU RAM and new generation CPU hardware.

Methodology notes

  • Industry/Industrial Analysis FrameworkUpstream, Midstream, and Downstream Industry Chain Transmission

    Software-Hardware Collaboration and Storage Hierarchy Transmission

    The report analyzes how changes in the AI software layer (such as long context, agent workflows) transmit downward to the hardware layer, leading to the evolution of storage architecture from a single HBM to a multi-level cache system of 'HBM-CPU RAM-SSD POD', helping readers understand how software demand reshapes the value distribution of the hardware industry chain.

  • Industry/Industrial Analysis FrameworkSupply and Demand Framework

    Reverse Constraint of Capacity Bottlenecks on Technical Specifications

    The report points out that Nvidia was forced to halve the memory capacity of the next-generation chip due to insufficient supplier LPDRAM capacity allocation, reflecting that in the semiconductor cycle, supply bottlenecks of upstream core components will directly reverse-constrain the technical specifications and iteration pace of downstream terminal products.

Asset mapping & comparison

Structured mapping from thesis to named assets (strengths, weaknesses, peers, risks).

  • Nvidia (NVDA)
    AI compute and architecture leader, launched Dynamo software, CMX platform, and Vera CPU, directly benefiting from the upgrade of software-hardware ecosystem in the inference era.
    Strengths
    Possesses a complete AI software-hardware ecosystem, capable of defining KV cache offloading standards and Agentic CPU architecture.
    Weaknesses
    Constrained by upstream LPDRAM capacity allocation, next-generation product specifications are forced to compromise.
    Comparison
    Leading traditional CPU vendors in defining the AI inference ecosystem.
    Risks
    Insufficient upstream storage chip capacity may limit the shipment and performance release of its high-end products.
  • Intel (INTC) / AMD (AMD)
    Traditional x86 CPU giants, benefiting from the rebound in CPU compute demand brought by Agentic AI and the reshaping of CPU/GPU ratios.
    Strengths
    Have deep accumulation and capacity foundation in the server CPU market.
    Weaknesses
    In a following state in the definition of AI-exclusive architecture.
    Comparison
    Facing fierce competition from Nvidia Vera and the Arm camp in the Agentic AI dedicated CPU market.
    Risks
    If failing to adapt to Agentic AI's demand for ultra-low latency architecture in time, may lose incremental market share.
  • SK hynix
    Core storage supplier, providing HBM, SOCAMM, and enterprise SSDs, directly benefiting from KV cache offloading and CPU RAM demand explosion.
    Strengths
    Leading layout in HBM and AI-related storage product lines, is a core supplier of Nvidia.
    Weaknesses
    LPDRAM capacity allocation faces multi-party gaming.
    Comparison
    In the first tier of the industry in terms of iteration speed of AI storage hardware.
    Risks
    If capacity expansion pace falls short of expectations, it will affect supply assurance for core customers.

Key data

  • Average Token Output Growth Rate per Question>5x/YearSurged since the second half of 2024, currently reaching approximately 30,000 to 40,000 tokens.
  • CPU-to-GPU Workload Ratio Evolution1:4/1:8 -> Approx. 1:1Agentic AI workflows lead to a significant rise in CPU compute demand.
  • Nvidia Vera CPU Memory Support1.5 TB LPDDR5XThree times the capacity of the previous generation Grace CPU.
  • CMX Platform Capacity per RackApprox. 9,600 TBManaged based on 64 BlueField-4 DPUs, used for Pod-level context storage.

Impact & implications

The report believes that the arrival of the AI inference era is completely revolutionizing storage systems. For the storage industry chain, SSD POD, as a new layer (G3.5) between local SSD and shared storage, will see continuous demand growth; for the semiconductor industry chain, Agentic AI's thirst for low latency and CPU compute not only opens up incremental space for CPU RAM but also makes 2026 an important turning point for the comprehensive refresh of CPU architecture across the industry to adapt to AI. Meanwhile, upstream LPDRAM capacity shortage has become a key bottleneck constraining AI chip specification upgrades.

Risks

  • Insufficient upstream LPDRAM capacity allocation may continue to constrain upgrades to AI chip and server memory specifications.
  • Slower-than-expected landing of AI inference applications and commercialization progress of Agentic AI may lead to slowed hardware incremental demand.

What to watch

  • Actual deployment scale and procurement pace of SSD POD platforms by giants such as Nvidia and Google.
  • Market acceptance and shipment volume of new Agentic AI products from major CPU vendors (Intel, AMD, Arm, etc.) in 2026.
  • Capacity expansion plans of upstream LPDRAM suppliers in 2027 and allocation adjustments for core customers such as Nvidia.
Zhejiang ICP No. 2022035445-5
Disclaimer: Market data, charts, indicators, research views, and other information provided on this website are intended solely for information display, research communication, and educational reference. They should not be regarded as personalized investment advice, securities recommendations, trading instructions, solicitations, or guarantees of return. While we strive to improve the reliability of our data and content, such information may still be subject to delays, errors, incompleteness, or untimely updates due to source differences, methodological limitations, system processing, or market volatility. Users should exercise independent judgment based on their own circumstances and bear all risks and responsibilities arising from the use of this website.

Settings

Sign in to view recent logins