Kioxia is working on XL-FLASH, a high-speed NAND tech designed to act as a DRAM alternative in AI servers. The goal: cut read latency to 1/10th of standard NAND, double capacity via CXL modules, and lift performance ~30%.
Their GP Series SSDs target 100M+ IOPS at 512-byte blocks for direct GPU access — helping offset scarce, expensive HBM. NVIDIA integration is slated for ~2028.
The real angle: AI inference and agentic apps need far more memory than HBM/DRAM can economically deliver. NAND is cheaper and scales better, so Kioxia is pushing it up the memory hierarchy.
If this works, it reshapes how AI infrastructure gets built — and who captures value in the stack.
Their GP Series SSDs target 100M+ IOPS at 512-byte blocks for direct GPU access — helping offset scarce, expensive HBM. NVIDIA integration is slated for ~2028.
The real angle: AI inference and agentic apps need far more memory than HBM/DRAM can economically deliver. NAND is cheaper and scales better, so Kioxia is pushing it up the memory hierarchy.
If this works, it reshapes how AI infrastructure gets built — and who captures value in the stack.
