vllm Releases
139 releases of vllm-project/vllm
- v0.10.0
The release initiates the cleanup of the V0 engine codebase by removing V0 CPU/XPU/TPU/HPU backends and various deprecated features while introducing new models and significant V1 engine improvements.
Jul 24, 2025
- release-proxy-8786
Minimal release notes: "Release PR https://github.com/neondatabase/neon/pull/12679. Diff with the previous release https://github.com/neondatab"
Jul 22, 2025
- release-compute-9011
Minimal release notes: "Release PR https://github.com/neondatabase/neon/pull/12653. Diff with the previous release https://github.com/neondatab"
Jul 18, 2025
- v0.9.2
The final V0 release featuring Priority Scheduling in V1, embedding models support, Mamba2 in V1, and full CUDA-Graph execution for FlashAttention v3 and FlashMLA.
Jul 7, 2025
- v0.9.1
Focuses on large-scale serving with DP/EP support, RLHF workflow improvements, Hybrid Memory Allocator, and FlexAttention integration into V1.
Jun 10, 2025
- v0.9.0.1
A patch release containing a critical bugfix for DeepSeek family models on NVIDIA Ampere and below.
May 30, 2025
- v0.3.0
This release introduces a multi-tier KVCache offloading framework and advanced routing strategies to optimize GPU memory usage and deployment density.
May 21, 2025
- v0.9.0
Major upgrade to PyTorch 2.7, initial set of optimized kernels on NVIDIA Blackwell, and initial DP/EP/PD support for large-scale inference.
May 15, 2025
- v0.8.5.post1
A post-release patch containing bugfixes for memory leaks and model accuracy issues.
May 2, 2025
- v0.8.5
Day 0 support for Qwen3, xgrammar structural tag feature, and important multi-modal bug fixes.
Apr 28, 2025
- v0.8.4
Includes important accuracy fixes for Llama4 models and adds support for Qwen3, Qwen3MoE, and xgrammar Enum support.
Apr 14, 2025
- v0.8.3
Day 0 Support for Llama 4 Scout and Maverick, V1 sliding window attention, and cluster scale serving features.
Apr 6, 2025
- v0.8.2
This release focuses on critical bug fixes for the V1 engine's memory usage and introduces new features like FP8 KV Cache support and the fastsafetensors loader.
Mar 23, 2025
- v0.8.1
A patch release addressing critical bug fixes for the V1 engine, including sampling parameter fixes and TPU support improvements.
Mar 19, 2025
- v0.8.0
A major release featuring the default enablement of the V1 engine, significant DeepSeek enhancements, and support for new models like Gemma 3 and Grok1.
Mar 18, 2025
- v0.0.3
Updates the rendering pipeline to support cameras with non-centered principal points and introduces texture mapping capabilities in FlameViewer.
Mar 7, 2025
- v0.7.3
This release highlights DeepSeek Multi-Token Prediction speedups and significant V1 engine features including LoRA, logprobs, and pipeline parallelism.
Feb 20, 2025
- v0.7.2
Introduces Qwen2.5-VL support, `transformers` backend support, and performance enhancements for DeepSeek models via MLA and FP8 kernels.
Feb 6, 2025
- v0.7.1
Focuses on MLA optimization for DeepSeek family models, V1 metrics and logging enhancements, and the addition of the MiniCPM-o model.
Feb 1, 2025
- v0.6.6.post1
A patch release that restores quantization support for other MoE methods which was broken by the initial DeepSeek V3 support.
Dec 27, 2024