Semiconductors and AI infrastructure Report Interpretation
JPMorgan finds that agentic-AI inference is shifting chip design toward memory capacity, bandwidth efficiency and low-latency interconnects. The report sees custom ASIC growth complementing rather than replacing merchant GPUs, while networking and HBM remain central constraints and opportunities.
Summary
JPMorgan finds that agentic-AI inference is shifting chip design toward memory capacity, bandwidth efficiency and low-latency interconnects. The report sees custom ASIC growth complementing rather than replacing merchant GPUs, while networking and HBM remain central constraints and opportunities.
- Decode-heavy agentic inference is making memory bandwidth, cache management and latency key architectural constraints.
- JPMorgan expects merchant GPUs and custom AI ASICs to coexist as complementary parts of the AI-compute stack.
- AMD's MI400 and Helios disclosures strengthened the report's view of its inference-share opportunity.
- Nvidia's Scale-In category and Broadcom's Thor Ultra highlight an expanding AI-networking opportunity.
- HBM3E, HBM4 and HBM4E availability remains a gating supply constraint for AI silicon.
Report Interpretation
Overview
This Hot Chips 2026 conference report argues that agentic-AI inference has become the semiconductor industry's central design imperative, broadening the AI infrastructure opportunity beyond training-focused accelerators. JPMorgan remains constructive on the multi-year buildout and identifies compute, custom ASICs, memory and networking beneficiaries.
Core views
JPMorgan’s central conclusion from Hot Chips 2026 is that agentic-AI inference has displaced training optimization as the principal chip-design focus. Multi-turn workloads, longer contexts, mixture-of-experts models, KV-cache management and low-latency multi-chip execution are making inference increasingly decode-dominated and memory-bandwidth-bound rather than primarily dependent on peak FLOPS. SambaNova quantified 75–97% of compute cycles for several frontier models as decode-dominated, which the report treats as evidence that memory-bandwidth utilization is becoming the binding constraint on inference throughput. Google’s inference-oriented TPU v8i allocates more on-chip SRAM for MoE latency, while Intel’s Crescent Island uses LPDDR5x to emphasize memory capacity. OpenAI likewise presented latency, tokens per joule and energy per request rather than peak FLOPS. JPMorgan therefore expects a growing split between training-optimized silicon, centered on peak compute and large scale-up bandwidth, and inference-optimized designs emphasizing capacity, SRAM and low-latency fabrics. Custom silicon remains a second major theme. OpenAI’s Jalapeno, its first internally designed inference ASIC developed with Broadcom, was presented as moving from initial RTL to tapeout in about nine months through AI-assisted hardware/software co-design. JPMorgan expects successor and adjacent programs to follow quickly. It places Jalapeno alongside Google TPUs, Meta MTIA, Microsoft MAIA, SambaNova’s SN50 RDU and Intel’s Crescent Island as evidence of a continuing internal-silicon push. Yet the report does not view ASICs as displacing merchant GPUs: custom designs fit scaled, well-defined workloads where customers can justify non-recurring engineering investment, while Nvidia and AMD platforms remain fungible options for the broader ecosystem. JPMorgan notes merchant GPUs currently hold more than 85% of AI-accelerator TAM, per JPMorgan estimates, but expects the merchant GPU versus custom XPU/ASIC mix to trend toward parity over the next several years. It also notes OpenAI’s stated plan for roughly 12GW of cumulative Nvidia infrastructure deployment through 2030 as evidence that custom and merchant capacity can coexist. AMD’s Instinct MI400 series and Helios rack were viewed as a more credible competitive response to Nvidia’s Vera Rubin. The Helios reference design provides 2.9 exaflops of AI compute and 31TB of HBM4 across 72 GPUs per rack, matching the reported 72-GPU Vera Rubin rack configuration. JPMorgan highlights MI455X’s chiplet architecture and stronger HBM4 capacity and bandwidth profile for MoE inference and long-context workloads. UALink128 is presented as an open-standard scale-up alternative to NVLink, while Helios combines GPUs, CPUs, DPUs and 800G AI NICs into a vertically integrated stack. Given the report’s view that inference demand is expanding rapidly, it sees AMD as positioned to gain AI-compute share from a low base, particularly in inference, though it does not provide a company rating in this report. Networking is becoming a broader and more differentiated part of AI infrastructure. Nvidia introduced BlueField-4’s “Scale-In Network” category for connecting GPU workloads with storage, CPUs running agents, orchestration and security, positioned between scale-up NVLink and scale-out Spectrum-X Ethernet. JPMorgan sees this as extending Nvidia’s networking stack across scale-up, scale-out, scale-across and scale-in. Broadcom’s Thor Ultra 800G Ethernet NIC, using Enhanced RoCE features to address reliability, congestion control and multipathing in AI clusters, supports the report’s view that Ethernet is gaining importance in scale-out and scale-across fabrics. The firm considers AI networking a rapidly growing TAM in which Broadcom and Nvidia are advancing with vertically integrated solutions, with spillover benefits for Marvell, Astera Labs and MTSI as cluster size and bandwidth needs rise. Broadcom is a cross-cutting beneficiary in JPMorgan’s analysis because of its involvement in Google TPUs, OpenAI’s Jalapeno, Meta MTIA, SambaNova’s RDU and Ethernet networking. The report sees this as expanding Broadcom’s role across both custom AI-compute silicon and AI infrastructure networking. HBM supply is the key common constraint: every major roadmap referenced dependence on HBM3E, HBM4 or HBM4E, reinforcing JPMorgan’s expectation of prolonged memory tightness. Intel’s Diamond Rapids Xeon and Crescent Island inference GPU show architectural progress, in JPMorgan’s view, but the report keeps execution risk elevated. Diamond Rapids’ 3D Mesh and disaggregated design are framed around hyperscale data movement, processing and security. Crescent Island is targeted at agentic inference with 480GB of LPDDR5x, a 350W TDP and on-chip MoE speculative-decoding acceleration. The report believes ultimate competitive impact depends on roadmap execution and customer commitments into 2027.
Analysis framework
JPMorgan synthesizes conference presentations to identify changes in AI workload characteristics, then links those changes to chip architecture, memory requirements, custom-silicon economics and networking design. It compares disclosed system specifications and design choices across vendors to assess likely beneficiaries and execution constraints.
Methodology notes
HBM supply-demand constraint
The report treats reliance on HBM3E, HBM4 and HBM4E across major AI roadmaps as evidence that memory supply remains a limiting factor for AI-silicon deployment.
AI workload changes transmitted through compute, memory and networking components
The report links agentic-inference workload requirements to accelerator architecture, HBM demand, interconnect fabrics, NICs, optical components and related semiconductor suppliers.
Asset mapping & comparison
Structured mapping from thesis to named assets (strengths, weaknesses, peers, risks).
- Broadcom (AVGO)Beneficiary of custom AI ASIC development and Ethernet-based AI networking.
- Strengths
- Involvement cited across Google TPU, OpenAI Jalapeno, Meta MTIA, SambaNova RDU and Thor Ultra Ethernet NICs.
- Comparison
- Positioned alongside Nvidia as a leading vertically integrated AI-networking incumbent.
- Risks
- HBM constraints may limit the pace of AI-silicon deployment.
- NVIDIA (NVDA)Merchant GPU and full-stack networking beneficiary.
- Strengths
- NVLink, Spectrum-X and BlueField-4 span scale-up, scale-out, scale-across and scale-in networking.
- Weaknesses
- Custom silicon is expected to gain share of AI-accelerator TAM over time.
- Comparison
- AMD’s Helios is positioned as a full-system alternative matching Nvidia’s cited 72-GPU rack configuration.
- Risks
- Custom ASIC expansion and HBM supply tightness.
- AMD (AMD)Potential AI-compute share gainer, especially in inference.
- Strengths
- MI400/Helios combines HBM4, UALink128 and integrated GPU, CPU, DPU and AI-NIC capabilities.
- Weaknesses
- The report describes share capture as coming from a low base.
- Comparison
- Presented as a competitive answer to Nvidia’s Vera Rubin and NVLink ecosystem.
- Risks
- Competitive outcomes depend on execution and ecosystem adoption.
- Marvell Technology (MRVL)Beneficiary of expanding AI scale-out networking and custom ASIC demand.
- Strengths
- Cited for custom scale-out silicon, DCI/DSPs, coherent optics and its Microsoft MAIA ASIC partnership.
- Comparison
- Benefits from Ethernet adoption alongside Broadcom.
- Astera Labs (ALAB)Beneficiary of greater cluster connectivity requirements.
- Strengths
- Cited for connectivity fabric and PCIe/CXL retimers.
- Comparison
- Part of the broader networking ecosystem rather than a merchant NIC/DPU incumbent.
- MTSIBeneficiary of rising AI-cluster bandwidth needs.
- Strengths
- Cited for high-speed lasers and optical components.
- Comparison
- Participates in the broader networking ecosystem supporting larger clusters.
- Micron Technology (MU)Beneficiary of sustained HBM demand.
- Strengths
- HBM remains a gating supply requirement across major AI-silicon roadmaps.
- SanDisk (SNDK)Named by JPMorgan as a beneficiary of the AI infrastructure buildout.
Key data
- Decode-dominated compute cycles75–97%SambaNova’s stated share across referenced frontier models.
- OpenAI Jalapeno development cycle~9 monthsFrom initial RTL to tapeout, according to the presentation.
- Merchant GPU share of AI accelerator TAM>85%Current share cited by JPMorgan; the report expects the mix with custom XPU/ASICs to move toward parity over several years.
- AMD Helios AI compute2.9 exaflopsAcross 72 GPUs per rack.
- AMD Helios HBM4 memory31TBRack-level memory capacity cited for the Helios reference architecture.
- Intel Crescent Island memory and power480GB LPDDR5x; 350W TDPSpecifications highlighted for its agentic-inference-oriented GPU.
Impact & implications
JPMorgan sees the shift toward agentic inference as broadening semiconductor demand beyond training-heavy accelerators to custom ASICs, memory, network interfaces, connectivity and optical components. It identifies AVGO, NVDA, AMD, MRVL, MTSI, ALAB, MU and SNDK as key beneficiaries of the continuing AI infrastructure buildout.
Risks
- HBM supply tightness may constrain AI-silicon availability and deployment.
- Intel’s competitive impact depends on roadmap execution and customer commitments into 2027.
What to watch
- The pace of custom XPU and ASIC deployments relative to merchant GPU demand.
- Availability of HBM3E, HBM4 and HBM4E across forthcoming AI-silicon roadmaps.
- AMD MI400 and Helios execution, UALink adoption and inference-market share gains.
- Further adoption of Ethernet and Enhanced RoCE in AI scale-out and scale-across networks.