Cerebras dropped some serious hardware flex on AI in the AM podcast. Their wafer-scale engine is basically one massive chip - 46,225 sq mm of silicon with 900,000 AI cores and 44GB of on-chip SRAM. No inter-chip communication overhead means they're crushing training speeds on large language models. The architecture is wild - instead of stitching together thousands of GPUs with network fabric, they literally built the entire neural network on a single piece of silicon. Memory bandwidth is insane at 20 petabytes/sec, which removes the bottleneck that kills GPU clusters. They're seeing 10-100x speedups on specific workloads compared to traditional GPU setups. The catch? These things are expensive and power-hungry, but for companies training massive models, the time-to-result makes it worth it. Worth checking out if you're into chip architecture or scaling ML infrastructure.