slime - GLM RL Training Framework
3.0slime is THUDM's open-source framework for LLM reinforcement learning, supporting GLM-4.5 and GLM-5.2. Integrates with Megatron and SGLang for Agentic RL.
About
slime is an open-source LLM post-training reinforcement learning framework developed by the THUDM team at Zhipu AI. It serves as the RL training infrastructure behind the GLM family of models, from GLM-4.5 through GLM-5.2. At its core, slime seamlessly connects the Megatron training engine with the SGLang inference framework, so training, inference, data generation, reward computation, and environment interaction all share a single unified pipeline. The key philosophy behind slime is to avoid turning the system into a heavy stack of separate trainers, inference services, and agent frameworks — instead, every component collaborates through a unified training/inference/data buffer path, keeping the system simple and scalable.
Key Features
Pricing & Fees
- Single unified pipeline that connects training, inference, data generation, reward computation, and environment interaction.
- Designed to avoid the complexity of stacking separate trainers, inference services, and agent frameworks.
- Direct parameter passthrough between Megatron and SGLang eliminates the need for additional abstraction layers.
- Combines BF16 training with FP8 inference and FP8 KV Cache to reduce memory overhead while maintaining model quality.
Who It's For
slime is built for LLM post-training — applying RL alignment training to pretrained models. It is also well suited for Agentic RL research, where developers explore reinforcement learning for tool calling and multi-step reasoning. Additionally, it supports model fine-tuning and compression workflows by leveraging the combination of FP8 inference and BF16 training to reduce resource requirements.