Over the past few days, the Hot Chips 2026 conference—lauded as a bellwether for the semiconductor industry—was held. Core vendors such as NVIDIA, Micron, and SK hynix delivered technical talks, focusing on the next evolution of AI compute infrastructure. In the past few years, companies have been competing on transistor density and peak compute performance. This year’s conference shows that the main bottleneck for AI compute has shifted from compute cores to data-movement efficiency and storage-system architecture. Memory is transitioning from being merely a data container into a system component. HBM is the core engine driving this new storage supercycle, but it faces physical constraints as well. Micron disclosed that, at the same capacity, the silicon die size of HBM3E is about three times that of traditional DDR5, due to its massive parallel data paths, TSVs, and power-delivery network. As a result, in typical GPU packaging, the memory chips account for about 90% of the silicon area at the system-level package. The enormous consumption of wafer capacity by HBM makes it difficult for the industry to quickly respond to demand shocks, leading to high pricing and adding a cost wall on top of the performance wall that data movement already imposes on AI systems. SK hynix also noted that as HBM progresses to 16 layers and even 20 layers, thermal dissipation and packaging are nearing physical limits, making the transition toward hybrid bonding an inevitable trend.

Facing bottlenecks, storage OEMs are beginning to extend toward logical computing. Samsung proposed a clear three-stage roadmap, aiming to turn HBM base die into an intelligent AI platform. First stage: customized HBM (cHBM). It transfers the memory controller to the HBM base die manufactured using advanced logic process technology. By using dedicated die-to-die (D2D) interfaces between chiplets in place of traditional HBM PHYs, it can reclaim 5% to 10% of silicon area for XPU use, translating into a 10% to 20% system performance improvement. Second stage: advanced HBM (AHBM). In underutilized regions of the base die, integrate processing units, offloading the attention mechanism—extremely bandwidth-intensive in large-model inference—to near-memory execution. Third stage: 3D-stacked HBM (zHBM). Abandon the 2.5D interposer layer and vertically stack HBM directly above the XPU. According to Samsung’s data, compared with standard HBM4E, zHBM’s I/O power is only about 0.5 pJ/bit, reducing total DRAM power by roughly 70%, boosting bandwidth by more than 2.3×. For a single module, it can save up to 100W—yielding about 8.3% of power headroom to be handed over to the accelerator. Achieving this requires wafer-to-wafer hybrid copper bonding with a pitch of less than 6 microns, and it also requires deep co-design from the start between DRAM and SoC logic.

At the inference edge, explosive capacity demand is driving a trend toward storage tiering. Given long contexts and large amounts of KV cache, simply relying on HBM is hard to justify on cost terms. Industry proposals for high-bandwidth flash (HBF) capacity—on the order of more than ten times that of HBM—offer strong economics: although per-GB bandwidth is only about 1/25 of HBM, in scenarios such as running trillion-parameter models on a single machine, low-batch processing, or sparse attention, HBF serves as an excellent high-capacity secondary memory layer. Increased demand for AI from mechanical hard drives has led to supply shortages. Large-scale data centers are starting to deploy lower-cost QLC SSDs more broadly between TLC SSDs and HDDs, improving system performance while also boosting NAND shipments.

For the market, Hot Chips’ release is not merely a short-term signal over a few quarters—it points to strategic trends over the next 3–5 years. The long-term value of the industry chain can be observed from three dimensions. First, storage OEMs: as HBM moves toward customization and even vertical integration with compute chips (zHBM), companies like Samsung and SK hynix are evolving from simply “selling memory” into joint designers of intelligent computing platforms. They are deeply involved in defining the architecture of next-generation accelerators, enabling stronger pricing power and greater inseparability. Second, advanced packaging and core equipment/materials: the evolution from micro-bumps to hybrid bonding, and from 2.5D to 3D WoW stacking, will bring significant incremental capital expenditures. This creates opportunities for upstream TSV etch equipment, hybrid bonding tools, testing equipment, and high-performance thermal interface materials. Domestic equipment/material suppliers that can be qualified simultaneously will benefit from scarce premium pricing. Third, general-purpose AI hardware: there will not be one-size-fits-all AI hardware. The training side will rely on zHBM that prioritizes bandwidth and high-density integration, while the inference side will give rise to a heterogeneous tiered system centered on HBF, high-capacity LPDDR, and enterprise-grade QLC SSDs. In the era of deep integration of compute-in-memory, the closer you are to the data, the greater the business value you create.