🤯 A 27B AI model is now running locally on an RTX 4060 with just 8GB VRAM.
Qwen 3.8-27B + Unsloth's new IQ4_XS quant:
→ 64K context
→ ~14.6GB on disk
→ ~150 tok/s prefill
→ ~5 tok/s decode
→ No GPU memory spillover
And Qwen 3.8-27B reportedly beats Claude Opus 4.6 on several benchmarks.
A frontier-level model running on a ~$300 consumer GPU.
Local AI is getting very serious. 🔥
Qwen 3.8-27B + Unsloth's new IQ4_XS quant:
→ 64K context
→ ~14.6GB on disk
→ ~150 tok/s prefill
→ ~5 tok/s decode
→ No GPU memory spillover
And Qwen 3.8-27B reportedly beats Claude Opus 4.6 on several benchmarks.
A frontier-level model running on a ~$300 consumer GPU.
Local AI is getting very serious. 🔥
