TokenSpeed: A Speed-of-Light LLM Inference Engine for Agentic Workloads
TokenSpeed is a high-performance LLM inference engine optimized for agentic workloads, featuring compiler-backed parallelism, a high-performance scheduler, and heterogeneous accelerator support.

Introduction
TokenSpeed is a speed-of-light LLM inference engine designed from first principles for agentic workloads. It incorporates a compiler-backed modeling mechanism for parallelism, a high-performance scheduler, safe KV resource reuse restriction, a pluggable layered kernel system that supports heterogeneous accelerators, and SMG integration.
Key Features
The tool is currently under continuous development. For more features, please visit the official website.
