Miles v0.1 provides a full-stack system for frontier-scale reinforcement learning, featuring rollout engines built on SGLang and backends for Megatron-LM and PyTorch FSDP. It supports full-parameter RL, LoRA RL, on-policy distillation, and supervised fine-tuning.
HOW THIS AFFECTS YOU
●
builderYou can deploy scalable, verified RL training loops using proven backends like Megatron-LM.
●
founderThis reduces the infrastructure complexity required to build frontier-level post-training pipelines.