Tool · Inference runtime · vLLM Project
vLLM
A high-throughput and memory-efficient inference and serving engine for LLMs
- License
- PermissiveApache License 2.0
- Language
- Python
- Latest release
- 21 Sept 2026vLLM v0.30.0
- Measurements
- 1with this runtime
Relationships
- ToolLiteLLMIntegrated by
Releases & news
- vLLM v0.30.0v0.30.0 Highlights This release features 762 commits from 315 contributors (104 new)! New models : DeepSeek-V4.1-Flash ( #56214 , #56228 , #56208 ) with the whole KV stored in MXFP8 through the FlashMLA V4.1 record on SM100 ( #56893 ), DeepGEMM Mega-mHC ( #56962 ), and async…
- vLLM proto-v0.3.0Release vllm-proto 0.3.0
- vLLM proto-v0.2.0Validated by PR #56538 CI at fa2a26f .
- proto-v0.1.0vllm-proto 0.1.0
- vLLM v0.29.0This release features 594 commits from 277 contributors (91 new)!
- vLLM v0.28.0This release features 584 commits from 270 contributors (76 new)!
- vLLM v0.27.1This is a patch release on top of v0.27.0.
- vLLM v0.27.0This release features 561 commits from 242 contributors (64 new)!
- vLLM v0.26.0This release features 411 commits from 212 contributors (61 new)!
- vLLM v0.25.1This release features 2 commits from 2 contributors (1 new)!
- vLLM v0.25.0This release features 558 commits from 232 contributors (64 new)!
- vLLM v0.24.0This release features 571 commits from 256 contributors (77 new)!
- vLLM v0.23.0Please note that Minimax M3 is not yet supported in this version. Please follow vLLM recipe for usage guides for M3.
- vLLM v0.22.1This release features 8 commits from 6 contributors (1 new)!
- vLLM V1 engine enters alphaRe-architected core engine with lower CPU overhead.
Measured performance with this runtime
Throughput members and projects recorded, in tokens per secondHow fast the model writes its answer; about 10 tokens per second reads like comfortable typing..
| Model · quantization | System | Runtime | Context | Measurements | Source |
|---|---|---|---|---|---|
| Qwen2.5 32B Instruct AWQ 4-bit | A100 80GB server (512 GB DDR4) | vLLMcuda · 0.7.3 | 8K | 48 tok/s · Generation throughput 72 GB · Peak memory 5,200 tok/s · Prompt throughput 95 ms · Time to first token | Mutinai illustrative fixtures |
Reviews
Contributions are not open yet, so there is nothing here from members.
Sources & history
Sources
- GitHub (live source, 3 records, 22 Sept 2026)
- GitHub (fixture) (illustrative fixture, 1 record, 13 Sept 2026)
- Mutinai illustrative fixtures (illustrative fixture, 1 record, 13 Sept 2026)
External identifiers
- github vllm-project/vllm
Field history
- homepageUrl = "https://vllm.ai" — GitHub (current)
- homepageUrl = "https://vllm.ai" — GitHub
- homepageUrl = "https://vllm.ai" — GitHub
- homepageUrl = "https://docs.vllm.ai" — GitHub (fixture)
- primaryLanguage = "Python" — GitHub (current)
- primaryLanguage = "Python" — GitHub
- primaryLanguage = "Python" — GitHub
- primaryLanguage = "Python" — GitHub (fixture)
- summary = "A high-throughput and memory-efficient inference and serving engine for LLMs" — GitHub (current)
- summary = "A high-throughput and memory-efficient inference and serving engine for LLMs" — GitHub
- summary = "A high-throughput and memory-efficient inference and serving engine for LLMs" — GitHub
- summary = "A high-throughput and memory-efficient inference and serving engine for LLMs" — GitHub (fixture)