SK Hynix just laid out the real AI bottleneck — and it's not compute anymore.

At the Dell forum, their system architecture VP said memory bandwidth now sets overall AI performance more than raw GPU power. LLMs already move massive data volumes. Agentic workloads and longer context windows push token throughput from tens per second toward hundreds or thousands.

HBM has scaled fast from Gen 1 through HBM4, but SK says it's no longer enough for next-gen agentic systems. Stacking memory directly on the GPU could unlock far higher bandwidth, but thermal limits make that path tough. Energy is the other pressure point — most power goes to moving data, not computing. That's why processing-in-memory research keeps advancing.

Memory pooling over CXL is being deployed to cut GPU idle time and improve cost efficiency. Efficiency gains elsewhere, including algorithmic improvements, have followed Jevons' paradox: they've increased total memory demand instead of reducing it.

The pattern is clear. Compute keeps advancing. The system is increasingly gated by how fast and how efficiently you can supply data. That's the constraint shaping the next wave of AI infrastructure spend.