Morgan Stanley’s latest research report says the AI memory shortage could persist throughout the entire cycle, but Nvidia won’t just wait for capacity to expand—it is reportedly preparing lower-spec versions of its Rubin platform, cutting HBM4 from 288GB to 192GB and LPDDR5 from 54TB to 28TB. Demand hasn’t disappeared; it is shifting toward NAND, DRAM, and high-speed interconnects. The CXL market could reach $6 billion by 2030. (Previously: Goldman Sachs raises its Kioxia price target to ¥116,000 and maintains a Buy rating, saying AI will keep NAND in short supply through 2028) (More background: YMTC plans to expand capacity by nearly 25%, says it will dominate NAND by 2027, and is reportedly planning a Q4 IPO to raise funds) While everyone is chasing GPU counts, Joseph Moore, a semiconductor analyst at Morgan Stanley, issued a counterintuitive warning in his latest report: the real bottleneck in the AI race has shifted to memory, and Nvidia’s solution is apparently to “cut specs.” The report reveals that Rubin’s HBM4 configuration could be reduced from 288GB to 192GB, while rack-level LPDDR5 could be cut in half, from 54TB to 28TB. This isn’t a sign that demand has peaked; rather, it marks the beginning of the entire AI industry being forced to redesign system architectures amid tight supply. TL;DR: The AI memory shortage could persist throughout the entire cycle. Morgan Stanley believes that continued growth in model size, context length, and concurrent inference workloads will keep absorbing new storage supply. Nvidia is beginning to offer “cut-spec” options for Rubin. Some HBM and LPDDR5 capacities may be reduced, but demand hasn’t disappeared—it is shifting toward NAND, DRAM, and high-speed interconnects. AI inference architectures are being restructured. Separating Prefill and Decode could improve hardware utilization and create opportunities for solutions from Cerebras, Nvidia Groq, and others. CXL could become a new growth driver. Morgan Stanley estimates that the related semiconductor market could reach around $6 billion by 2030, with Astera Labs and Marvell as potential beneficiaries. For storage stocks, the key is how long the cycle lasts, not the size of short-term price increases. Morgan Stanley maintains Overweight ratings on Micron and SanDisk. The main risks are a slowdown in AI investment and data center construction. The following is a translation of the original article: AI infrastructure is facing an increasingly thorny problem: computing power can be expanded quickly by buying more GPUs, but memory supply cannot keep pace. Morgan Stanley believes that AI demand will continue to absorb substantial amounts of new memory supply for some time. Although the severity of shortages may fluctuate over time, DRAM capacity expansion is unlikely to close the supply-demand gap quickly. After speaking with several computing companies, Morgan Stanley found that finding ways around the memory bottleneck has become an increasingly important topic. Nvidia CEO Jensen Huang has also discussed the need to ease supply constraints through the co-design of computing, networking, and storage systems. Analysts believe this is not a warning that storage demand is weakening, but a reality the entire AI industry must face: when memory cannot be supplied as planned, companies have to find new system architectures that allow existing hardware to continue supporting AI growth. The most straightforward response is to lower the memory specifications built into individual servers or GPUs. Nvidia’s Rubin to “Cut Specs”: HBM4 Reduced to 192GB Morgan Stanley notes that tight DRAM and NAND supply, along with rising prices, is forcing the AI supply chain to reconsider product configurations. For chipmakers, rather than sticking to the original specifications and risking systems not being delivered as planned, it makes more sense to offer versions with different memory capacities so customers can choose according to their needs. The report expects Nvidia may offer several lower-memory configurations for the Rubin platform. For rack-level memory, Rubin was originally planned to have around 54TB of LPDDR5, but some new configurations may bring this down to 28TB, with SOCAMM2 memory module capacity falling from 192GB to 96GB. For HBM, each Rubin GPU was originally expected to come with 288GB of HBM4, using eight 12-high stacks. Morgan Stanley expects Nvidia may add a product version with 192GB of HBM4, reducing the stack height to eight layers. For Rubin Ultra, the report considers 512GB a more appropriate point of comparison following a shift to a dual-compute-die design. Different capacity options of around 192GB to 384GB may emerge in the future. These are Morgan Stanley’s assessments of product configurations as of the report’s publication date; Nvidia has not officially confirmed all of the specifications. From a cost perspective, the strategy makes sense. Reducing the number of HBM stack layers lowers capacity, while theoretical bandwidth does not necessarily fall in tandem if the number of interfaces and pin speeds remain unchanged. But capacity and bandwidth address two different problems. For applications with smaller models and shorter contexts, lower capacity may have limited impact. But as models grow and inference tasks become more complex, insufficient memory capacity may mean more data has to move between storage tiers, increasing latency or GPU utilization. More importantly, reducing HBM does not make the data AI needs to process disappear; it simply means that data has to find a new place to reside. For example, when GPU HBM capacity is insufficient, some data may need to be stored in main memory. When rack-level LPDDR5 capacity is reduced, some KV Cache (key-value cache) requirements may shift to NAND flash memory. KV Cache is used to store historical computation data during large language model inference. Its capacity requirements rise as contexts grow longer and more requests run concurrently. Memory demand hasn’t disappeared; it is simply moving to other tiers. Morgan Stanley notes that while Nvidia is reducing some LPDDR5 configurations, the industry has already observed additional demand for NAND. Lower HBM capacity could also increase data transfers between GPUs, putting more pressure on high-speed interconnect networks. Demand Hasn’t Disappeared: Data Is Being Pushed Toward NAND and Interconnect Chips This means Nvidia’s cut-spec strategy may temporarily ease DRAM supply constraints, while simultaneously increasing demand for NAND, interconnect chips, and larger GPU clusters. There are long-term structural factors behind this shift in pressure. According to historical data from Epoch AI cited in the report, the parameter count of frontier models in the large language model era has grown at a rate of roughly doubling every six months. Context windows have also expanded rapidly. At the same time, more AI applications are moving beyond simple question-and-answer tasks toward programming, complex reasoning, and long-running agent tasks, further increasing memory usage. Morgan Stanley believes that model size, context length, and inference concurrency will continue to drive storage demand. Therefore, reducing the specifications of some products is more likely to be a temporary solution during a period of tight supply than a sign of a long-term decline in AI memory demand. If reducing memory configurations is a short-term response, changing how AI inference is computed could bring deeper changes to the industry. The second path Morgan Stanley is watching... “Why Is the AI Memory Shortage Prompting Nvidia to Cut Specs? Morgan Stanley: Bottlenecks Will Push Data Toward NAND and Interconnect Chips” was first published on BlockTempo (the most influential blockchain news media outlet).
