Tool · Inference runtime · turboderp
ExLlamaV2
A fast inference library for running LLMs locally on modern consumer-class GPUs
- License
- PermissiveMIT License
- Language
- Python
- Latest release
- 13 Jul 2025ExLlamaV2 0.3.2
- Measurements
- 0with this runtime
Relationships
No relationships recorded.
Releases & news
- ExLlamaV2 0.3.0Add Qwen3 and Qwen3MoE support
- ExLlamaV2 0.2.9Add Torch 2.7.0 wheels (big thanks to @kingbri1 for unborking the build action)
- ExLlamaV2 0.2.8Support Qwen2.5-VL
- ExLlamaV2 0.2.7Basic video support for Qwen2-VL
- ExLlamaV2 0.2.6Some small fixes, most notably for Qwen2-VL inference on Windows
- ExLlamaV2 0.2.5Initial support for Qwen2-VL (images for now, no video)
- ExLlamaV2 0.2.4Support Pixtral
- ExLlamaV2 0.2.3No longer use safetensors for loading weights (fix virtual memory issues on Windows especially)
Measured performance with this runtime
Throughput members and projects recorded, in tokens per secondHow fast the model writes its answer; about 10 tokens per second reads like comfortable typing..
No performance measurements recorded yet.
Reviews
Contributions are not open yet, so there is nothing here from members.
Sources & history
Sources
- GitHub (live source, 2 records, 13 Sept 2026)
External identifiers
- github turboderp/exllamav2
Field history
- homepageUrl = null — GitHub (current)
- homepageUrl = null — GitHub
- primaryLanguage = "Python" — GitHub (current)
- primaryLanguage = "Python" — GitHub
- summary = "A fast inference library for running LLMs locally on modern consumer-class GPUs" — GitHub (current)
- summary = "A fast inference library for running LLMs locally on modern consumer-class GPUs" — GitHub