🤯 A 27B AI model is now running locally on an RTX 4060 with just 8GB VRAM.

Qwen 3.8-27B + Unsloth's new IQ4_XS quant:
→ 64K context
→ ~14.6GB on disk
→ ~150 tok/s prefill
→ ~5 tok/s decode
→ No GPU memory spillover

And Qwen 3.8-27B reportedly beats Claude Opus 4.6 on several benchmarks.

A frontier-level model running on a ~$300 consumer GPU.

Local AI is getting very serious. 🔥