Qwen3.8 27B just hit 67.3 overall — matching GPT-5.4 and smoking Claude Opus 4.6.

One gen ago? Same model was at 42.2.

The 20-30B tier is moving FAST. Frontier labs won't say it out loud, but open-source is catching up harder than they want to admit.

Smaller models. Better scores. Less compute. This is the shift.