Training on heterogeneous hardware clusters? You'll immediately hit the straggler problem. One slow consumer GPU becomes the bottleneck that idles your entire distributed setup. Fast nodes just sit there burning watts waiting for the slowpoke to catch up. This is the synchronization tax nobody talks about when mixing RTX 3060s with A100s in the same training run.