vllm Releases
139 releases of vllm-project/vllm
- v0.2.0Meetily v0.2.0
This major update brings custom OpenAI integration, streamlined onboarding, automatic updates, and enhanced reliability to Meetily, a meeting minutes application with comprehensive bug fixes.
Dec 31, 2025
- v0.13.0
Focuses on model support expansion, engine core improvements like conditional compilation and prefix caching, and breaking changes related to PassConfig and environment variables.
Dec 19, 2025
- v0.6.0v0.6
This release introduces static site publishing and port endpoints to Apptron, a browser-based Linux development environment. It also enhances WASM support and improves custom environment build performance.
Dec 6, 2025
- v0.12.0
Introduces GPU Model Runner V2, significant speculative decoding improvements, and prepares for PyTorch 2.9.0 upgrade.
Dec 3, 2025
- v0.11.2
A patch release fixing critical bugs related to Ray, speculative decoding, async-scheduling, and SM100 CUTLASS MoE macro.
Nov 20, 2025
- v0.11.1
A large release updating to PyTorch 2.9.0, enabling batch-invariant torch.compile, and stabilizing async scheduling.
Nov 18, 2025
- v0.5.0v0.5
This release introduces basic project visibility and sharing capabilities to Apptron, allowing users to keep projects private or share them via URL, along with performance improvements and Safari support.
Nov 7, 2025
- v0.4.0v0.4
This is the first official release of Apptron, deployed to apptron.dev. It offers basic project creation and a Wanix-powered in-browser development environment with cloud sync capabilities.
Oct 24, 2025
- 0.1.1🎉 Meetily v0.1.1 - Major Release
A major architectural overhaul transforming the application into a standalone app with significant performance improvements, new recording features, and enhanced AI summarization capabilities.
Oct 23, 2025
- v0.11.0
Marks the end of V0 engine support and turns on FULL_AND_PIECEWISE CUDA graph mode by default.
Oct 2, 2025
- v0.10.2
Adds native aarch64 support, new models, and hardware optimizations for Blackwell/SM100 generation.
Sep 13, 2025
- v0.16.7Move flash imports
Minimal release notes: "## What's Changed * Move flash attention funcs by @VikParuchuri in https://github.com/datalab-to/surya/pull/457 **"
Sep 8, 2025
- v0.16.6Enable setting attention method
This release focuses on configuration flexibility by removing strict checks on the attention method and allowing users to set it manually.
Sep 8, 2025
- v6.5.0Rustlings 6.5.0
Upgrades the Rust toolchain to edition 2024 and raises the minimum supported Rust version, along with improvements to the file watcher and VS Code integration.
Aug 21, 2025
- v0.10.1.1
A critical patch release addressing security vulnerabilities and a critical bug in CUTLASS MLA Full CUDAGraph.
Aug 20, 2025
- v0.4.1
A maintenance release containing cherry-picked bug fixes and KVCache integration updates from various pull requests for the release-0.4 branch.
Aug 19, 2025
- v0.10.1
A major release focusing on V0 engine deprecation, extensive new model support including GPT-OSS and vision-language models, and significant hardware optimizations for NVIDIA Blackwell and AMD ROCm platforms.
Aug 18, 2025
- release-proxy-8853
Minimal release notes: "Release PR https://github.com/neondatabase/neon/pull/12766. Diff with the previous release https://github.com/neondatab"
Jul 29, 2025
- release-compute-9073
Minimal release notes: "Release PR https://github.com/neondatabase/neon/pull/12738. Diff with the previous release https://github.com/neondatab"
Jul 28, 2025
- release-9129
Minimal release notes: "Release PR https://github.com/neondatabase/neon/pull/12737. Diff with the previous release https://github.com/neondatab"
Jul 25, 2025