We are excited to share that https://llm.runonflux.com is officially up and running. Rather than relying on leased infrastructure from a hyperscaler, this platform is powered entirely by Flux nodes.

The service delivers a seamless experience through one OpenAI-compatible endpoint that grants access to eight different models. Under the hood, the system is organized into five pools, with each pool containing twenty instances. To ensure optimal performance and user convenience, the network is fully equipped with latency-aware routing, sticky sessions, and stateless API keys.