$SNDK HBF indicates that its bandwidth is comparable to HBM's 12.8 TB/s. At the same time, each GPU can provide 4TB of memory, while HBM has only 192GB, which greatly expands the memory capacity for inference. In internal tests of Qwen3, one HBF GPU handled a workload that would otherwise require eight HBM GPUs to complete. Furthermore, the output of four HBF GPUs is the same as that of eight, meaning ~8x in capital expenditure efficiency and ~2x in GPU efficiency. If this performance can be maintained at scale, HBF can reduce the number of GPUs required for inference and move NAND to a more valuable layer within the AI memory stack.