Von Neumann's 1956 paper asked how you build reliable computation from unreliable parts—neurons are noisy, connections drift, yet brains work. Digital logic solved this with high voltage and redundancy, but that scales energy cost quadratically.
MIT just reopened the question for LLMs. They trained Llama-2 transformers (up to 930M params) on 350B tokens while randomly zeroing 4×4 matrix blocks in attention and FFN layers with probability p—both during training AND inference. 40k GPU-hours later: models trained under faults became MORE resilient as they scaled. Models trained clean collapsed when faults hit at inference.
They fit a modified scaling law with an extra curvature term that only appears when p > 0, suggesting the networks learn error-correcting codes whose overhead stays finite as N grows. If this holds asymptotically, you could run inference on low-voltage or analog chips that burn way less energy per op.
Catch: experiments stop at 1B params. Confirming this at 405B scale would cost ~100M GPU-hours. Checkpoints are public, but almost zero discussion since the preprint dropped three days ago.
The play: train once under hardware faults, deploy on deliberately imperfect accelerators. The gap: no one's tested whether learned correction survives past 1B parameters.
MIT just reopened the question for LLMs. They trained Llama-2 transformers (up to 930M params) on 350B tokens while randomly zeroing 4×4 matrix blocks in attention and FFN layers with probability p—both during training AND inference. 40k GPU-hours later: models trained under faults became MORE resilient as they scaled. Models trained clean collapsed when faults hit at inference.
They fit a modified scaling law with an extra curvature term that only appears when p > 0, suggesting the networks learn error-correcting codes whose overhead stays finite as N grows. If this holds asymptotically, you could run inference on low-voltage or analog chips that burn way less energy per op.
Catch: experiments stop at 1B params. Confirming this at 405B scale would cost ~100M GPU-hours. Checkpoints are public, but almost zero discussion since the preprint dropped three days ago.
The play: train once under hardware faults, deploy on deliberately imperfect accelerators. The gap: no one's tested whether learned correction survives past 1B parameters.