Report Interpretation
Covering the latest research from top Wall Street investment banks
Report InterpretationHilo Research

Report Interpretation

The report proposes a vertical-die V-die architecture and validates its feasibility using models spanning electromagnetic, system, and thermal scales. The results show a 4.01-fold increase in bandwidth, an 81.2% improvement in LLM decoding speed, a 2.2-fold improvement in throughput stability under context scaling, and a 43.7% reduction in peak temperature compared with 16-Hi HBM4.

IndustrySemiconductor Memory and Advanced Packaging

Summary

The V-die 3.5D architecture overcomes the I/O, capacity, and cooling bottlenecks of conventional HBM through sidewall interconnects and direct cooling

The report proposes a vertical-die V-die architecture and validates its feasibility using models spanning electromagnetic, system, and thermal scales. The results show a 4.01-fold increase in bandwidth, an 81.2% improvement in LLM decoding speed, a 2.2-fold improvement in throughput stability under context scaling, and a 43.7% reduction in peak temperature compared with 16-Hi HBM4.

—
High-Bandwidth MemoryV-die3.5D IntegrationAdvanced PackagingArtificial IntelligenceMemory WallDirect Liquid CoolingChiplet Systems
  • V-die vertically positions DRAM dies and connects them to bottom bumps through the sidewalls, reducing dependence on the number of through-silicon vias.
  • The configuration comparison shows that the pin count increases from 2048 to 8192, while theoretical memory bandwidth rises from 10.24 TB/s to 40.96 TB/s.
  • Although the 9 mm RDL routing is longer than HBM's 5 mm routing, it still meets the JEDEC HBM4 receiver eye-mask standard at 8 Gbps.
  • The LLM simulation is summarized as delivering an 81.2% improvement in decoding speed and a 2.2-fold improvement in throughput stability as context length scales.
  • TSV-free V-die DRAM can save 25% in power consumption and reclaim die area.
  • Coolant flows directly through the gaps between dies, reducing peak temperature by 43.7% compared with 16-Hi HBM4.

Report Interpretation

Overview

This is a semiconductor architecture study addressing the AI memory-wall problem. The report argues that the conventional HBM scaling path based on per-channel data rate, stack height, and I/O width is increasingly constrained by TSV area, package height, and heat accumulation. The proposed V-die 3.5D integration provides greater I/O and thermal scalability through sidewall I/O, bottom bumps, and direct inter-die cooling.

Core views

The report first attributes the problem to the mismatch between the growth of AI model size and the expansion rate of memory capacity per GPU: the parameter count of advanced Transformer models is described as increasing 410-fold every two years, while memory per GPU increases only twofold. Historically, HBM has scaled primarily along three dimensions: per-channel data rate, stack height, and I/O width. The generational path presented in the report spans HBM1 (2014), HBM2 (2018), HBM2E (2020), HBM3 (2022), HBM3E (2024), and HBM4 (2026), with stacks progressing from 4-Hi toward 12-Hi/16-Hi, per-channel rates rising from 1 Gbps to 8 Gbps, and I/O counts increasing from 1024 to 2048. However, this scaling approach is increasingly subject to physical constraints. The first constraint is that TSVs occupy usable DRAM area, preventing their number from increasing indefinitely. The second is package height along the Z-axis: the table indicates typical total heights of approximately 720, 720, and 775 micrometers for HBM1/HBM2, HBM3, and HBM4, respectively, while the typical per-layer thicknesses for 8-Hi, 12-Hi, and 16-Hi configurations are approximately 90, 60, and 48.4 micrometers. Adding more layers under a fixed height constraint requires thinner dies, which can affect yield and increase thermal density. The third constraint is the vertical thermal path: heat from the bottom dies must pass through multiple layers of BEOL, microbumps, underfill, and potential microvoids. Thermal resistance accumulates layer by layer, creating a pronounced temperature gradient and cooling bottleneck. The V-die design arranges DRAM dies vertically and implements interconnections through sidewall I/O pads and bottom bumps. This structure no longer leaves I/O scaling entirely constrained by TSVs within the die, while the gaps between dies increase the surface area accessible to coolant. In a 16-Hi configuration, the report argues that this layout offers stronger I/O scalability, but the trade-off is that the RDL must extend to every die. The longest routing path is approximately 9 mm, compared with approximately 5 mm for conventional HBM, making it necessary to verify whether the additional insertion loss, reflections, and crosstalk remain acceptable. Electrical validation calibrates electromagnetic simulations using S-parameter measurements from a 40 GHz, 50-ohm matched coplanar waveguide. The extracted relative permittivity is 6.5, the loss tangent is 0.002, and copper conductivity is 3×10^7 S/m. The RDL is modeled as copper signal traces embedded in silicon dioxide between upper and lower ground planes. The results show that, at the maximum rate of 8 Gbps, the passive-channel characteristics of V-die are generally comparable to those of HBM. The longer traces increase insertion loss, but return loss remains similar, while near-end and far-end crosstalk both remain below −30 dB across the entire frequency range. At 8 Gbps, the V-die channel satisfies the JEDEC HBM4 receiver eye-mask requirements of 0.3 UI and 100 mV, although its eye margin is slightly lower than that of HBM. The system-level analysis incorporates the electromagnetically validated parameters into a JEDEC-timing-based memory model and uses gem5 for synthetic-traffic simulation. Both comparison configurations use 16 memory stacks and an 8.0 Gbps pin rate. The configuration table shows that the number of pins per stack increases from 2048 to 8192, memory bandwidth rises from 10.24 TB/s to 40.96 TB/s, and L2 cache bandwidth correspondingly increases from 24.0 TB/s to 96.0 TB/s to prevent the cache side from masking memory I/O capability. Based on these results, the report ultimately concludes that V-die can deliver 4.01 times higher bandwidth. At the application level, LLMCompass is used to evaluate the A100 model and a newly constructed H100 model for this study. The workload is a 175-billion-parameter model with 96 layers, 96 attention heads, and a hidden dimension of 12288. The report measures inference decoding performance using token generation rate and operation-level latency, concluding that V-die delivers an 81.2% improvement in decoding speed. In the context-scaling test, the results page indicates a 24.8% speed reduction, while the report concludes that V-die provides 2.2 times greater throughput stability than the reference design, demonstrating that higher and more stable memory bandwidth can alleviate memory bottlenecks in long-context inference. The power and thermal analysis further explains V-die's scaling advantages. The HBM4 base die consumes 9 W, comprising 6 W for the PHY and 3 W for TSVs. Each HBM4 DRAM die consumes approximately 2 W, including TSV overhead. V-die DRAM eliminates TSVs, which the report states can save 25% in power consumption while freeing die area. Conventional HBM4 relies on a top cold plate, requiring heat to travel through the entire stack. V-die instead allows coolant to flow directly through the gaps between dies, enabling heat generated by each die to be absorbed locally; the inter-die coolant effectively acts as a thermal ground plane. Under identical conditions with an ambient temperature of 45°C, V-die's peak temperature is lower than HBM4's, eliminating the pronounced temperature gradient between dies. The final result is a 43.7% reduction in peak temperature compared with 16-Hi HBM4. The report therefore concludes that V-die simultaneously mitigates I/O, capacity, and thermal scaling constraints through vertical routing and direct cooling, and that joint modeling across electromagnetic, system, application, and thermal scales is critical for determining whether the architecture can translate into improved LLM performance. The study also emphasizes that system simulation in the era of chiplets and heterogeneous integration must incorporate cross-domain physical constraints rather than infer application performance solely from nominal bandwidth.

Analysis framework

The study proceeds in the sequence of problem definition, architecture design, physical validation, system simulation, application testing, and thermal validation. It first uses generational HBM specifications to explain the constraints imposed by TSVs, package height, and thermal paths, then proposes a V-die design featuring sidewall interconnects and inter-die cooling. It subsequently evaluates RDL signal integrity using a measurement-calibrated electromagnetic model, feeds the validated parameters into a gem5 synthetic-traffic model and an LLMCompass workload model, and finally compares the thermal performance of HBM4 and V-die using power maps and thermal-resistance networks.

Methodology notes

  • (Method Outside the Taxonomy)

    Cross-Scale Joint Validation Framework

    The report progressively propagates physical-layer electromagnetic and thermal constraints into system and LLM application models to determine whether architectural changes can produce actual improvements in bandwidth, latency, and decoding throughput.

  • (Method Outside the Taxonomy)

    Electromagnetic Simulation Calibrated by S-Parameter Measurements

    The study uses 40 GHz, 50-ohm coplanar-waveguide measurements to extract permittivity, loss, and copper conductivity, then uses these parameters to simulate the insertion loss, reflections, crosstalk, and eye diagrams of the longer RDL traces.

  • (Method Outside the Taxonomy)

    Synthetic-Traffic and LLM Workload Simulation

    The study first uses gem5 and a JEDEC timing model to test bandwidth and latency under traffic loads, then uses LLMCompass to evaluate the token generation rate and operation-level latency of a 175-billion-parameter model.

  • (Method Outside the Taxonomy)

    Thermal-Resistance Network Analysis

    The report compares the heat-conduction paths of a top cold plate and direct inter-die cooling, analyzing total thermal resistance, peak temperature, and temperature gradients across different die layers.

Key data

  • Model Parameter Growth vs. GPU Memory Scaling410-fold every two years; 2-fold for memory per GPUUsed by the report to illustrate the growth mismatch underlying the AI system memory wall
  • HBM Generational SpanHBM1 (2014) to HBM4 (2026)Stacks scale from 4-Hi to 12-Hi/16-Hi, while per-channel rates rise from 1 Gbps to 8 Gbps
  • Typical Per-Layer Die Thickness90 micrometers, 60 micrometers, 48.4 micrometersCorresponding to 8-Hi, 12-Hi, and 16-Hi configurations, respectively, illustrating the pressure to thin dies under a fixed package-height constraint
  • RDL Path LengthV-die 9 mm, HBM 5 mmV-die's longer routing increases insertion loss, but return loss remains similar
  • Electromagnetic Model Calibration40 GHz, 50 ohmsMaterial parameters are extracted by correlating coplanar-waveguide S-parameter measurements with simulations
  • Extracted Material Parametersεr=6.5, tanδ=0.002, copper conductivity=3×10^7 S/mUsed for high-fidelity RDL routing simulation
  • Crosstalk LevelBoth NEXT and FEXT below −30 dBAcross the entire analyzed frequency range
  • Eye-Diagram Compliance8 Gbps; 0.3 UI, 100 mVV-die meets the JEDEC HBM4 receiver eye mask, but its margin is slightly lower than HBM's
  • Pins per StackIncreased from 2048 to 8192Comparison configuration for HBM4 and V-die
  • Memory BandwidthIncreased from 10.24 TB/s to 40.96 TB/sConfiguration values; the report concludes that actual bandwidth improves by 4.01 times
  • L2 Cache BandwidthIncreased from 24.0 TB/s to 96.0 TB/sRaised to prevent the cache side from masking memory I/O capability
  • LLM Workload175 billion parameters, 96 layers, 96 attention heads, hidden dimension of 12288Used to simulate token generation rate and operation-level latency
  • LLM Decoding Performance81.2% fasterThe report's summary of V-die's relative performance
  • Context-Scaling Stability2.2 timesThe report's summarized improvement in throughput stability; the results page also indicates a 24.8% speed reduction
  • HBM4 Base-Die Power Consumption9 W6 W for the PHY and 3 W for TSVs
  • Per-Die HBM4 DRAM Power ConsumptionApproximately 2 WIncludes TSV overhead
  • V-die DRAM Power Savings25%Resulting from the elimination of TSVs, while also reclaiming die area
  • Peak Temperature ImprovementReduced by 43.7%Compared with 16-Hi HBM4 under the same ambient temperature of 45°C

Impact & implications

The report argues that if sidewall I/O and direct inter-die cooling can be implemented as validated, V-die could bypass the conventional HBM scaling path that depends on more TSVs and thinner dies, allowing higher capacity and wider I/O without incurring equivalent increases in area, power consumption, and thermal density. For AI systems, these improvements could translate into higher LLM decoding throughput and greater performance stability for long-context workloads. The study also demonstrates that novel chiplet architectures must be validated using joint models spanning electromagnetic, thermal, and application layers.

Risks

  • V-die must use RDL connections up to approximately 9 mm long to connect dies at each layer. Compared with HBM's approximately 5 mm paths, this results in higher insertion loss and slightly reduced eye-diagram margin.
  • The architecture involves tightly coupled electrical, timing, capacity, and thermal constraints. The report explicitly states that cross-domain evaluation is necessary and that nominal performance at a single layer is insufficient to demonstrate system feasibility.

What to watch

  • Further technological progress in interconnects, manufacturing, and cooling structures along the V-die packaging roadmap.
  • Whether the cross-scale validation framework can be extended to chiplet-based heterogeneous integration systems and reusable LLM workload modules.

Settings

Sign in to view recent logins