v0.14.1

vllm-project/vllmv0.14.1Jan 24, 2026by khluu

AI Summary

A patch release focused on security vulnerabilities and memory leak fixes following v0.14.0.

Key Highlights

  • Security fixes for HTTP header limits and eval usage
  • Memory leak fixes
  • Bugfixes for async-scheduling and FlashAttn MLA

New Features

  • Bugfix for Ray with multiple nodes
  • Fix false assertion with spec-decode and TP>2
  • Fix async-scheduling + FlashAttn MLA compatibility

Full Release Notes

This is a patch release on top of `v0.14.0` to address a few security and memory leak fixes.