NVIDIA advances Rubin Ultra and multi-rack scaling; Kyber commercialization may lag
AI summary card
NVIDIA advances Rubin Ultra and multi-rack scaling; Kyber commercialization may lag
Vera Rubin Strata entered production at the end of July, but Rubin Ultra's chips, HBM, and rack solutions remain under iteration; NVIDIA may first use multiple Oberon racks in H2 2027 to meet surging compute demand.
- Vera Rubin Strata entered commercial production at the end of July after a brief delay, with enhanced Vera CPUs and Rubin GPUs expected by the end of Q4 2026.
- Dual-chiplet and quad-chiplet Rubin Ultra configurations are being developed in parallel; the dual-chiplet version may fit the existing Oberon NVL72, while the quad-chiplet version targets Kyber systems.
- Multiple Oberon racks can scale through copper or optical interconnects to 144, 288, or 576 GPUs, and the research team expects prototype hardware to arrive as early as H2 2027.
- The HBM selection for Rubin Ultra remains undecided; training places greater emphasis on bandwidth, while long-context agent inference depends more on memory capacity and high-speed external KV cache.
- The bill of materials for dual-GPU Rubin Ultra Strata is expected to increase by about 30% versus Rubin Strata in H2 2026; if gross margins remain at 75% to 80%, ASP could exceed $170,000.
Report interpretation
Overview
This report updates the chip- and rack-level development progress of NVIDIA's Vera Rubin and Rubin Ultra platforms. The central view is that Vera Rubin has entered production and continues to undergo performance optimization, while Rubin Ultra's design and system architecture have yet to be finalized; given Kyber's high complexity, NVIDIA may first use a multi-Oberon-rack solution to bridge the product-evolution gap.
Core views
For Vera Rubin, the research team expects NVIDIA to continue optimizing the Vera CPU and Rubin GPU through tape-out iterations, targeting peak TDP of about 2,300W and claimed FP4 inference performance of 50PF, about 3.3 times that of GB300; FP4 performance for the standard 1,800W SKU may not exceed 40PF. For Rubin Ultra, the dual-GPU version may retain the Oberon NVL72 form factor, while the quad-GPU version corresponds to Kyber; the quad-chiplet SKU has not been cancelled, but commercialization may be slower. Multi-rack systems can expand a unified GPU pool to as many as 576 GPUs and, together with networking and storage racks, improve bandwidth, compute capability, and cost competitiveness per token.
Analysis framework
The report conducts scenario analysis of chip topology, HBM specifications, rack interconnects, system bandwidth, and pricing based on industry research, product-roadmap inferences, public demonstration examples, and bill-of-materials estimates.
Methodology notes
Evaluates the Vera CPU, Rubin GPU, HBM, packaging, and rack systems in layers.
By comparing dual-chiplet, quad-chiplet, Oberon, Kyber, and multi-rack solutions, it assesses performance-upgrade paths and commercialization timing.
Derives module selling prices from component costs, memory share, and target gross margins.
The research team expects Rubin Ultra Strata's cost increase to be driven mainly by HBM, SoCAAM, logic devices, and packaging-related components, and uses this to derive potential ASP.
Training workloads emphasize memory bandwidth, while long-context agent inference emphasizes capacity and KV-cache storage.
The report argues that simply raising HBM bandwidth cannot resolve insufficient memory capacity, and that CMX/eSSD storage architectures can supplement long-context inference.
Asset mapping & comparison
Structured mapping from thesis to named assets (strengths, weaknesses, peers, risks).
- NVDA.USCore covered name and platform supplier
- Strengths
- Vera Rubin is already in mass production; CPU/GPU performance is being continuously optimized; the company has system-level product capabilities spanning chips, packaging, networking, and storage racks; multi-rack architectures can improve total bandwidth and flexibility.
- Weaknesses
- Dual-GPU Rubin Ultra performance has not yet been disclosed; HBM capacity and specifications have not been finalized; costs and power consumption for high-end systems continue to rise.
- Comparison
- Compared with a single NVL72 rack, multi-rack systems can provide higher aggregate bandwidth; compared with the complex Kyber, expansion solutions based on the existing Oberon may be deployed faster.
- Risks
- Delays in tape-out and system integration, postponed Kyber commercialization, HBM supply and cost pressure, unclear customer acceptance of an approximately 30% price increase, and accelerated in-house ASIC development by customers.
Key data
- Vera Rubin Strata mass-production timingEnd of July 2026The report states that it entered commercial production after a brief delay.
- Enhanced Vera Rubin peak TDPApproximately 2,300WThe target level expected by the research team for the enhanced CPU and GPU combination.
- Vera Rubin FP4 inference performance50PFThis is the company's claimed target; the report says it is about 3.3 times that of GB300.
- Expected Rubin Ultra tape-outQ4 2026The report's expectation for GPU chiplet tape-out; the specific timing remains uncertain.
- Rubin Ultra peak TDPApproximately 2,600WApplies to the Vera Rubin Ultra dual-GPU configuration estimated by the research team.
- Rubin Ultra quad-chiplet performance target100PF FP4At GTC 2026, NVIDIA mentioned a quad-GPU version paired with 1T HBM4E; dual-GPU performance was not disclosed.
- Multi-rack GPU scale144 to 576 GPUsCan scale through two-rack, four-rack, or eight-rack configurations.
- Potential multi-rack deployment timingH2 2027Based on industry research inferences regarding prototype hardware paired with Rubin Ultra.
- HBM4E bandwidth per stackUp to approximately 4.1TBpsBased on an example with a 16Gbps speed and 2,048-bit bus width.
- Total bandwidth of 12 HBM4E stacksApproximately 49.2TBpsThe report estimates this is 50% higher than eight HBM modules.
- Rubin Ultra Strata cost changeApproximately +30%Relative to the forecast for Rubin Strata in H2 2026.
- Potential Rubin Ultra Strata ASPOver $170,000Assumes NVIDIA maintains gross margins of 75% to 80%.
Impact & implications
For NVIDIA, multi-rack scaling can sustain compute supply and system-level differentiation while Kyber advances, while supporting different AI workloads through high-bandwidth interconnects and CMX storage. For cloud service providers and hyperscale customers, improved system capability may reduce cost per token, but total capital expenditure will also rise further from 2027 to 2028, potentially driving external financing needs or accelerating in-house ASIC development.
Risks
- Rubin Ultra chips and rack architectures remain under development, and tape-out, validation, and mass-production timelines may be delayed.
- Kyber's quad-chiplet GPU and rack-system design are complex, and commercialization may proceed more slowly than multi-Oberon solutions.
- HBM4E capacity, speed, and bandwidth configurations have not been finalized; lower capacity may constrain long-context agent inference.
- Price increases for HBM, SoCAAM, packaging, and other components may raise bill-of-materials costs and reduce pricing flexibility.
- If the dual-GPU Rubin Ultra's performance gain is insufficient, customers may not accept a potential price increase of approximately 30%.
- Multi-rack systems will increase customer capital expenditure, potentially prompting hyperscale customers to accelerate in-house ASIC development to reduce dependence on commercial GPUs.
What to watch
- Tape-out, performance, and supply progress for enhanced Vera CPUs and Rubin GPUs in Q4 2026.
- Official specifications, FP4 performance, and HBM4E configurations for dual-chiplet and quad-chiplet Rubin Ultra SKUs.
- Product finalization and commercialization timing for Kyber racks, as well as deployment of multi-Oberon/Polyphe systems in H2 2027.
- Deployment progress for CMX and BlueField 4 STX storage racks and eSSD supplier-certification results.
- Changes in HBM, SoCAAM, and packaging costs, as well as Rubin Ultra pricing and customer procurement feedback.
- Capital expenditure, financing activity, and in-house ASIC strategies of cloud service providers and hyperscale customers.