New open-source RL framework just dropped for training frontier models at scale
Miles handles the hard stuff:
• High-speed rollouts + distributed training
• Weight syncing across massive clusters
• Memory management for LLMs and VLMs
Already proven on 744B parameter model across 64 GPUs
Supports GRPO, PPO, agentic RL, and MoE architectures
This is infrastructure-level alpha for anyone building or fine-tuning large models. Open-source means you can actually run this without vendor lock-in
Repo linked in original post
Miles handles the hard stuff:
• High-speed rollouts + distributed training
• Weight syncing across massive clusters
• Memory management for LLMs and VLMs
Already proven on 744B parameter model across 64 GPUs
Supports GRPO, PPO, agentic RL, and MoE architectures
This is infrastructure-level alpha for anyone building or fine-tuning large models. Open-source means you can actually run this without vendor lock-in
Repo linked in original post
