Open source LLMs are scaling hard. GLM-5.3 hit 320B params, Kimi K3 pushed to 2.8T, Llama 4 runs 109B active (likely MoE architecture).

Devs aren't just doing inference anymore—they're fine-tuning on proprietary datasets, spinning up domain-specific variants, and deploying production systems on open weights. The shift is real: open models are now viable for serious workloads, not just experiments.