Tools
6 of 13ToolWorks withLicenseMeasurements
Runtime · turboderp · Python
A fast inference library for running LLMs locally on modern consumer-class GPUs
Works withexl2NVIDIA GPUs
LicensePermissive
Measurements—
Runtime · ggml.org · C++
LLM inference in C/C++
Works withggufNVIDIA GPUs, AMD GPUs, Apple silicon, most GPUs (Vulkan), CPUs
LicensePermissive
Runtime · Apple · Python
Run LLMs with MLX
Works withmlxApple silicon
LicensePermissive
Runtime · Ollama · Go
Get up and running with Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
Works withggufNVIDIA GPUs, AMD GPUs, Apple silicon, CPUs
LicensePermissive
Measurements—
Runtime · SGLang Project · Python
SGLang is a high-performance serving framework for large language models and multimodal models.
Works withsafetensorsNVIDIA GPUs, AMD GPUs
LicensePermissive
Measurements—
Runtime · vLLM Project · Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Works withsafetensorsNVIDIA GPUs, AMD GPUs
LicensePermissive