Public preview · figures are illustrative fixture data, attributed to their source · data to May 2025 · read-only
Tool · Inference runtime · turboderp

ExLlamaV2

A fast inference library for running LLMs locally on modern consumer-class GPUs

License
PermissiveMIT License
Language
Python
Latest release
13 Jul 2025ExLlamaV2 0.3.2
Measurements
0with this runtime

Relationships

No relationships recorded.

Releases & news

  1. ExLlamaV2 0.3.0
    Add Qwen3 and Qwen3MoE support
  2. ExLlamaV2 0.2.9
    Add Torch 2.7.0 wheels (big thanks to @kingbri1 for unborking the build action)
  3. ExLlamaV2 0.2.8
    Support Qwen2.5-VL
  4. ExLlamaV2 0.2.7
    Basic video support for Qwen2-VL
  5. ExLlamaV2 0.2.6
    Some small fixes, most notably for Qwen2-VL inference on Windows
  6. ExLlamaV2 0.2.5
    Initial support for Qwen2-VL (images for now, no video)
  7. ExLlamaV2 0.2.4
    Support Pixtral
  8. ExLlamaV2 0.2.3
    No longer use safetensors for loading weights (fix virtual memory issues on Windows especially)

Measured performance with this runtime

Throughput members and projects recorded, in tokens per secondHow fast the model writes its answer; about 10 tokens per second reads like comfortable typing..

No performance measurements recorded yet.

Reviews

Contributions are not open yet, so there is nothing here from members.

Sources & history

Sources

  • GitHub (live source, 2 records, 13 Sept 2026)

External identifiers

Field history

  • homepageUrl = null GitHub (current)
  • homepageUrl = null GitHub
  • primaryLanguage = "Python" GitHub (current)
  • primaryLanguage = "Python" GitHub
  • summary = "A fast inference library for running LLMs locally on modern consumer-class GPUs" GitHub (current)
  • summary = "A fast inference library for running LLMs locally on modern consumer-class GPUs" GitHub