AI Inference Reshapes Memory Demand: KV Cache Offloading and Agentic AI Become Core Drivers
AI summary card
AI Inference Reshapes Memory Demand: KV Cache Offloading and Agentic AI Become Core Drivers
Hardware demand in the AI inference era is shifting from training to inference; KV cache offloading technology spawns new demand for SSD PODs, while the explosion of Agentic AI pushes the CPU-to-GPU compute ratio towards 1:1, significantly driving CPU RAM demand.
- AI inference imposes three new requirements on hardware: higher QPS, longer context windows, and more inference steps.
- Token output per question surges over 5x, marking the industry's entry into the Test-time Scaling phase.
- KV cache occupancy explodes with context and Batch Size; Nvidia launches Dynamo software and CMX platform to offload KV cache to CPU RAM and SSD POD.
- Agentic AI requires significant CPU for orchestration and tool invocation; the CPU-to-GPU workload ratio will shift from 1:4 or 1:8 to 1:1.
- Nvidia Vera CPU supports 1.5TB LPDDR5X, but limited by LPDRAM capacity, the next-generation module memory capacity is forced to halve.
- 2026 will see a comprehensive refresh of Agentic AI-exclusive CPUs, with Intel, AMD, Arm, and Ampere all launching new products.
Report interpretation
Overview
This research report, published by TrendForce, deeply analyzes the structural transformation brought to memory and storage systems in the AI inference era. The report points out that as AI moves from training to inference, the core contradiction of hardware demand is shifting. The explosive growth of KV Cache has spawned new demand for storage hierarchy offloading, while the rise of Agentic AI has completely reshaped the compute ratio between CPU and GPU, thereby opening up a huge incremental market for CPU RAM and enterprise SSDs.
Core views
Hardware demand in the AI inference phase differs fundamentally from the training phase, mainly reflected in higher Queries Per Second (QPS), longer context windows, and more inference steps. Nvidia data shows that since the second half of 2024, the average token output per question has surged at a rate of over 5x per year, reaching 30,000 to 40,000 tokens, marking the industry's entry into the 'Test-time Scaling' thinking phase, which directly translates into huge consumption of memory and compute. In terms of memory occupancy, model weights belong to static allocation, while KV cache is dynamically allocated. As conversation length and Batch Size increase, the memory consumption of KV cache rises sharply. When GPU HBM capacity is insufficient, the system has to discard KV cache and recalculate, leading to increased latency and higher Total Cost of Ownership (TCO). To this end, Nvidia launched Dynamo software to offload infrequently accessed KV cache to CPU RAM and SSDs, which have lower bandwidth but larger capacity and better cost efficiency. Meanwhile, Nvidia released the CMX Context Memory Storage platform in early 2026, adding a Pod-level context layer (SSD POD, i.e., G3.5 layer) between local SSD and shared storage via BlueField-4 DPU, specifically designed to manage the massive KV cache generated by long-context workloads. On the other hand, the explosion of Agentic AI is reshaping CPU architecture. Agents need to actively plan, invoke tools, and execute operations, which requires the CPU to undertake tasks such as orchestration, data routing, and sub-agent evaluation. Therefore, the CPU-to-GPU workload ratio is evolving from the traditional 1:4 or 1:8 towards 1:1. Nvidia launched the Vera CPU designed specifically for agents, supporting up to 1.5TB of LPDDR5X memory, three times that of the previous generation Grace. However, due to insufficient supplier LPDRAM capacity allocation in 2027, Nvidia was forced to halve the SOCAMM memory capacity of the next-generation Vera Rubin superchip module. Besides Nvidia, Intel, AMD, Arm, and Ampere are also launching new products intensively in 2026, welcoming the comprehensive CPU replacement wave brought by Agentic AI.
Analysis framework
The report adopts an analytical framework of 'Demand Breakdown and Architecture Evolution'. First, the institution breaks down the hardware demand of AI inference into two dimensions: model weights (static) and KV cache (dynamic). By quantifying the surge trend of Token output, it deduces the HBM capacity bottleneck, thereby引出 (leading to) the technical path of 'KV cache offloading' and its driving effect on SSD POD and CPU RAM. Second, starting from the workflow characteristics of Agentic AI (planning, tool invocation), the institution deduces its dependence on low latency and CPU compute, thus concluding the industry trend of CPU/GPU ratio converging to 1:1, and finally landing on the incremental demand for CPU RAM and new generation CPU hardware.
Methodology notes
Software-Hardware Collaboration and Storage Hierarchy Transmission
The report analyzes how changes in the AI software layer (such as long context, agent workflows) transmit downward to the hardware layer, leading to the evolution of storage architecture from a single HBM to a multi-level cache system of 'HBM-CPU RAM-SSD POD', helping readers understand how software demand reshapes the value distribution of the hardware industry chain.
Reverse Constraint of Capacity Bottlenecks on Technical Specifications
The report points out that Nvidia was forced to halve the memory capacity of the next-generation chip due to insufficient supplier LPDRAM capacity allocation, reflecting that in the semiconductor cycle, supply bottlenecks of upstream core components will directly reverse-constrain the technical specifications and iteration pace of downstream terminal products.
Asset mapping & comparison
Structured mapping from thesis to named assets (strengths, weaknesses, peers, risks).
- Nvidia (NVDA)AI compute and architecture leader, launched Dynamo software, CMX platform, and Vera CPU, directly benefiting from the upgrade of software-hardware ecosystem in the inference era.
- Strengths
- Possesses a complete AI software-hardware ecosystem, capable of defining KV cache offloading standards and Agentic CPU architecture.
- Weaknesses
- Constrained by upstream LPDRAM capacity allocation, next-generation product specifications are forced to compromise.
- Comparison
- Leading traditional CPU vendors in defining the AI inference ecosystem.
- Risks
- Insufficient upstream storage chip capacity may limit the shipment and performance release of its high-end products.
- Intel (INTC) / AMD (AMD)Traditional x86 CPU giants, benefiting from the rebound in CPU compute demand brought by Agentic AI and the reshaping of CPU/GPU ratios.
- Strengths
- Have deep accumulation and capacity foundation in the server CPU market.
- Weaknesses
- In a following state in the definition of AI-exclusive architecture.
- Comparison
- Facing fierce competition from Nvidia Vera and the Arm camp in the Agentic AI dedicated CPU market.
- Risks
- If failing to adapt to Agentic AI's demand for ultra-low latency architecture in time, may lose incremental market share.
- SK hynixCore storage supplier, providing HBM, SOCAMM, and enterprise SSDs, directly benefiting from KV cache offloading and CPU RAM demand explosion.
- Strengths
- Leading layout in HBM and AI-related storage product lines, is a core supplier of Nvidia.
- Weaknesses
- LPDRAM capacity allocation faces multi-party gaming.
- Comparison
- In the first tier of the industry in terms of iteration speed of AI storage hardware.
- Risks
- If capacity expansion pace falls short of expectations, it will affect supply assurance for core customers.
Key data
- Average Token Output Growth Rate per Question>5x/YearSurged since the second half of 2024, currently reaching approximately 30,000 to 40,000 tokens.
- CPU-to-GPU Workload Ratio Evolution1:4/1:8 -> Approx. 1:1Agentic AI workflows lead to a significant rise in CPU compute demand.
- Nvidia Vera CPU Memory Support1.5 TB LPDDR5XThree times the capacity of the previous generation Grace CPU.
- CMX Platform Capacity per RackApprox. 9,600 TBManaged based on 64 BlueField-4 DPUs, used for Pod-level context storage.
Impact & implications
The report believes that the arrival of the AI inference era is completely revolutionizing storage systems. For the storage industry chain, SSD POD, as a new layer (G3.5) between local SSD and shared storage, will see continuous demand growth; for the semiconductor industry chain, Agentic AI's thirst for low latency and CPU compute not only opens up incremental space for CPU RAM but also makes 2026 an important turning point for the comprehensive refresh of CPU architecture across the industry to adapt to AI. Meanwhile, upstream LPDRAM capacity shortage has become a key bottleneck constraining AI chip specification upgrades.
Risks
- Insufficient upstream LPDRAM capacity allocation may continue to constrain upgrades to AI chip and server memory specifications.
- Slower-than-expected landing of AI inference applications and commercialization progress of Agentic AI may lead to slowed hardware incremental demand.
What to watch
- Actual deployment scale and procurement pace of SSD POD platforms by giants such as Nvidia and Google.
- Market acceptance and shipment volume of new Agentic AI products from major CPU vendors (Intel, AMD, Arm, etc.) in 2026.
- Capacity expansion plans of upstream LPDRAM suppliers in 2027 and allocation adjustments for core customers such as Nvidia.