TokenSpeed: A Speed-of-Light LLM Inference Engine for Agentic Workloads

TokenSpeed: A Speed-of-Light LLM Inference Engine for Agentic Workloads

TokenSpeed is a high-performance LLM inference engine optimized for agentic workloads, featuring compiler-backed parallelism, a high-performance scheduler, and heterogeneous accelerator support.

TokenSpeed: A Speed-of-Light LLM Inference Engine for Agentic Workloads screenshot

Introduction

TokenSpeed is a speed-of-light LLM inference engine designed from first principles for agentic workloads. It incorporates a compiler-backed modeling mechanism for parallelism, a high-performance scheduler, safe KV resource reuse restriction, a pluggable layered kernel system that supports heterogeneous accelerators, and SMG integration.

Key Features

The tool is currently under continuous development. For more features, please visit the official website.

Scroll to Top