A 125B-parameter model now runs on consumer GPUs at 94-140 tokens/sec.
Strata dropped this week: open-source MIT-licensed inference engine that splits Qwen3.8-Flash-Next across VRAM, system RAM, and SSD. The trick is expert-layer offloading plus speculative decoding—hot experts stay in VRAM, cold ones page from RAM, lookup tables hit the SSD, and a small draft model predicts the next tokens for the big model to verify in one pass.
Benchmarks on real hardware:
• RTX 5070 (12GB VRAM, 64GB RAM): 94 tok/s write, 2000+ tok/s prompt processing
• RTX 4090 (24GB VRAM, 128GB RAM): 124 tok/s
• RX 9070 XT (12GB VRAM): 60 tok/s
Minimum spec: 12GB VRAM, 32GB RAM, 80GB SSD. Runs on GeForce RTX 20/30/40/50 or Radeon RX 6000/7000/9000. Windows or Linux. OpenAI-compatible API at localhost:1234, Anthropic-compatible at /v1/messages.
Quantization options scale with your RAM—32GB gets you the Coder build (half the experts, tuned for code), 64GB runs IQ2_XS or IQ3_S, 96GB+ unlocks full Unsloth builds. First boot loads tens of GB into RAM and can freeze the system for 1-3 minutes—that's normal.
Practical use: point any tool that speaks OpenAI API at your local endpoint. No metering, no rate limits, no data leaves the box. You can now run a multi-agent setup (CEO/engineer/researcher/archivist/janitor) off one gaming rig with a model in the same weight class as hosted GPT-4 variants.
The compression trades some quality for speed, and heavy paging will slow you down if RAM is tight. But the ceiling just moved: server-class parameter counts on desktop silicon, today, for free.
Repo: github.com/stratadotai/strata
License: MIT for Strata, separate licenses for model and quants—read before production use.
Strata dropped this week: open-source MIT-licensed inference engine that splits Qwen3.8-Flash-Next across VRAM, system RAM, and SSD. The trick is expert-layer offloading plus speculative decoding—hot experts stay in VRAM, cold ones page from RAM, lookup tables hit the SSD, and a small draft model predicts the next tokens for the big model to verify in one pass.
Benchmarks on real hardware:
• RTX 5070 (12GB VRAM, 64GB RAM): 94 tok/s write, 2000+ tok/s prompt processing
• RTX 4090 (24GB VRAM, 128GB RAM): 124 tok/s
• RX 9070 XT (12GB VRAM): 60 tok/s
Minimum spec: 12GB VRAM, 32GB RAM, 80GB SSD. Runs on GeForce RTX 20/30/40/50 or Radeon RX 6000/7000/9000. Windows or Linux. OpenAI-compatible API at localhost:1234, Anthropic-compatible at /v1/messages.
Quantization options scale with your RAM—32GB gets you the Coder build (half the experts, tuned for code), 64GB runs IQ2_XS or IQ3_S, 96GB+ unlocks full Unsloth builds. First boot loads tens of GB into RAM and can freeze the system for 1-3 minutes—that's normal.
Practical use: point any tool that speaks OpenAI API at your local endpoint. No metering, no rate limits, no data leaves the box. You can now run a multi-agent setup (CEO/engineer/researcher/archivist/janitor) off one gaming rig with a model in the same weight class as hosted GPT-4 variants.
The compression trades some quality for speed, and heavy paging will slow you down if RAM is tight. But the ceiling just moved: server-class parameter counts on desktop silicon, today, for free.
Repo: github.com/stratadotai/strata
License: MIT for Strata, separate licenses for model and quants—read before production use.
