vllm Releases
139 releases of vllm-project/vllm
- v0.21.0
A major release deprecating Transformers v4, requiring C++20, and integrating KV Offload + HMA.
May 15, 2026
- v0.20.2
This is a small patch release that addresses critical stability issues in DeepSeek V4, gpt-oss, and Qwen3-VL models, specifically fixing hangs, KV cache allocation failures, and kernel capture problems.
May 10, 2026
- v0.20.1
Patch release focused on DeepSeek V4 stabilization and performance improvements, along with several important bug fixes. Includes multi-stream pre-attention GEMM, BF16/MXFP8 all-to-all support, and integrated tile kernels for optimized head computation.
May 4, 2026
- v0.20.0
Major release with 752 commits from 320 contributors featuring CUDA 13.0 default, PyTorch 2.11 upgrade, Transformers v5 support, DeepSeek V4 initial support, FlashAttention 4 as default MLA prefill, and TurboQuant 2-bit KV cache. Includes vLLM IR skeleton and significant Model Runner V2 advances.
Apr 27, 2026
- v0.72.00.72.0
This release introduces a new 'inline' style for headers and footers to embed them within the list frame, alongside a new 'dashed' border style. It also improves Vim integration by handling window resizing events and fixes several bugs related to UI display and keyboard input.
Apr 26, 2026
- v0.19.1
Patch release on top of v0.19.0 with Transformers v5.5.3 upgrade and comprehensive bug fixes for Gemma4 including streaming tool calls, HTML duplication, and token repetition issues.
Apr 18, 2026
- v0.71.00.71.0
Focuses on shell integration improvements, cross-reload item identity tracking, and performance scaling improvements with CPU cores.
Apr 4, 2026
- v0.19.0
Major release with 448 commits from 197 contributors featuring full Gemma 4 architecture support, zero-bubble async scheduling with speculative decoding, Model Runner V2 maturation, ViT Full CUDA Graphs, and general CPU KV cache offloading.
Apr 3, 2026
- v0.18.1
Patch release addressing specific issues including reverting SM100 MLA prefill backend, fixing mock.patch resolution failure for Python <= 3.10, and addressing DeepGemm E8M0 accuracy degradation for Qwen3.5 FP8.
Mar 31, 2026
- v0.18.0
A major feature release introducing gRPC serving, GPU-less rendering, and NGram GPU speculative decoding. It includes significant hardware support for AMD ROCm and Intel XPU, along with improvements to KV cache offloading and Elastic Expert Parallelism.
Mar 20, 2026
- v0.17.1
A patch release addressing issues with TRTLLM fused MoE, DeepSeek V3.2, and adding support for the Nemotron 3 Super model.
Mar 11, 2026
- v0.17.0
A major upgrade to PyTorch 2.10 and integration of FlashAttention 4. Introduces Model Runner V2 with Pipeline Parallel and Decode Context Parallel support, alongside full Qwen3.5 model family support.
Mar 7, 2026
- v0.7.0v0.7
Focuses on web-based environment features including console mode embeds, offline site caching, and performance improvements for network throughput and environment build times.
Mar 3, 2026
- v0.70.00.70.0
Focuses on dynamic option binding and performance optimizations for the fzf CLI tool.
Mar 2, 2026
- v0.16.0
A major release featuring Async scheduling and Pipeline Parallelism for significant throughput improvements, a new Realtime API for streaming audio, and a major overhaul of the XPU platform.
Feb 25, 2026
- v0.2.1Meetily v0.2.1
A maintenance release for the Meetily application fixing Whisper model handling, Windows audio device selection issues, and API connection problems.
Feb 11, 2026
- v0.15.1
A patch release focusing on security fixes, RTX Blackwell GPU support corrections, and various bug fixes to improve stability and performance.
Feb 4, 2026
- v0.15.0
This release features 335 commits from 158 contributors, introducing new model architectures like Kimi-K2.5 and Molmo2, enhancing engine core capabilities with async scheduling and pipeline parallelism, and optimizing performance for NVIDIA Blackwell and AMD ROCm platforms.
Jan 29, 2026
- v0.14.1
A patch release focused on security vulnerabilities and memory leak fixes following v0.14.0.
Jan 24, 2026
- v0.14.0
A major update enabling async scheduling by default, requiring PyTorch 2.9.1, and adding features like gRPC server, `--max-model-len auto`, and Model Runner V2.
Jan 20, 2026