Public preview · figures are illustrative fixture data, attributed to their source · data to May 2025 · read-only
Tool · Inference runtime · vLLM Project

vLLM

A high-throughput and memory-efficient inference and serving engine for LLMs

License
PermissiveApache License 2.0
Language
Python
Latest release
21 Sept 2026vLLM v0.30.0
Measurements
1with this runtime

Relationships

Releases & news

  1. vLLM v0.30.0
    v0.30.0 Highlights This release features 762 commits from 315 contributors (104 new)! New models : DeepSeek-V4.1-Flash ( #56214 , #56228 , #56208 ) with the whole KV stored in MXFP8 through the FlashMLA V4.1 record on SM100 ( #56893 ), DeepGEMM Mega-mHC ( #56962 ), and async…
  2. vLLM proto-v0.3.0
    Release vllm-proto 0.3.0
  3. vLLM proto-v0.2.0
    Validated by PR #56538 CI at fa2a26f .
  4. proto-v0.1.0
    vllm-proto 0.1.0
  5. vLLM v0.29.0
    This release features 594 commits from 277 contributors (91 new)!
  6. vLLM v0.28.0
    This release features 584 commits from 270 contributors (76 new)!
  7. vLLM v0.27.1
    This is a patch release on top of v0.27.0.
  8. vLLM v0.27.0
    This release features 561 commits from 242 contributors (64 new)!
  9. vLLM v0.26.0
    This release features 411 commits from 212 contributors (61 new)!
  10. vLLM v0.25.1
    This release features 2 commits from 2 contributors (1 new)!
  11. vLLM v0.25.0
    This release features 558 commits from 232 contributors (64 new)!
  12. vLLM v0.24.0
    This release features 571 commits from 256 contributors (77 new)!
  13. vLLM v0.23.0
    Please note that Minimax M3 is not yet supported in this version. Please follow vLLM recipe for usage guides for M3.
  14. vLLM v0.22.1
    This release features 8 commits from 6 contributors (1 new)!
  15. vLLM V1 engine enters alpha
    Re-architected core engine with lower CPU overhead.

Measured performance with this runtime

Throughput members and projects recorded, in tokens per secondHow fast the model writes its answer; about 10 tokens per second reads like comfortable typing..

Model · quantizationSystemRuntimeContextMeasurementsSource
Qwen2.5 32B Instruct AWQ 4-bitA100 80GB server (512 GB DDR4)vLLMcuda · 0.7.38K
48 tok/s · Generation throughput
72 GB · Peak memory
5,200 tok/s · Prompt throughput
95 ms · Time to first token
Mutinai illustrative fixtures

Reviews

Contributions are not open yet, so there is nothing here from members.

Sources & history

Sources

  • GitHub (live source, 3 records, 22 Sept 2026)
  • GitHub (fixture) (illustrative fixture, 1 record, 13 Sept 2026)
  • Mutinai illustrative fixtures (illustrative fixture, 1 record, 13 Sept 2026)

External identifiers

Field history

  • homepageUrl = "https://vllm.ai" GitHub (current)
  • homepageUrl = "https://vllm.ai" GitHub
  • homepageUrl = "https://vllm.ai" GitHub
  • homepageUrl = "https://docs.vllm.ai" GitHub (fixture)
  • primaryLanguage = "Python" GitHub (current)
  • primaryLanguage = "Python" GitHub
  • primaryLanguage = "Python" GitHub
  • primaryLanguage = "Python" GitHub (fixture)
  • summary = "A high-throughput and memory-efficient inference and serving engine for LLMs" GitHub (current)
  • summary = "A high-throughput and memory-efficient inference and serving engine for LLMs" GitHub
  • summary = "A high-throughput and memory-efficient inference and serving engine for LLMs" GitHub
  • summary = "A high-throughput and memory-efficient inference and serving engine for LLMs" GitHub (fixture)