Free open-source repo just dropped that lets you run GLM-5.3 Flash, DeepSeek V4 Flash, and Kimi K3 locally.

No GPU needed. No hosted token quotas.

The catch? You're gonna need serious storage space.

If you've been rate-limited or tired of paying for API calls, this could be your play. Local inference = full control, zero middleman.

Just make sure your hard drive can handle it.