vllm-omni Releases
48 releases of vllm-project/vllm-omni
- mcp/v1.0.2
Introduces the Hook0 MCP server for AI assistant integration alongside dependency updates.
May 11, 2026
- play/v1.0.3
Adds a full CLI tool, web UI for webhook testing, and SEO improvements to the Hook0 Play environment.
May 10, 2026
- output-worker/v1.0.2
Implements a robust output worker with Pulsar support, bounded retries, and object storage capabilities.
May 10, 2026
- v3.6.11
Enhances OneNote import functionality and resolves issues with the Rich Text Editor in secondary windows.
May 8, 2026
- v0.20.0
v0.20.0 is a major production release aligned with upstream vLLM v0.20.0. It refreshes the serving/runtime stack for large-scale omni workloads, expands model coverage across speech, omni, image, audio, and video generation, and improves performance, quantization, and hardware readiness across CUDA, ROCm, MUSA, NPU, and XPU backends.
May 7, 2026
- v1.39.0
A major security and feature update for Logto focusing on key rotation, authentication customization, and new connectors.
Apr 30, 2026
- v2.2.0v2.2.0-beta — Inbound Voice AI Infrastructure
This release focuses on infrastructure changes to make inbound voice AI production-ready, featuring distributed SIP registration and improved multi-server failover capabilities.
Apr 28, 2026
- v2.1.0v2.1.0 — Built-In Observability
This release introduces built-in observability, a refactored streaming pipeline architecture, and improved developer experience for the voice AI platform.
Apr 1, 2026
- v1.38.0
This release focuses on enhancing authentication capabilities, specifically targeting device flow, passkey support, and comprehensive session management.
Mar 31, 2026
- v0.18.0
v0.18.0 is a major rebase and systems release aligning with upstream vLLM v0.18.0. It features a refactored serving entrypoint architecture, strengthened audio/speech production serving, substantial diffusion optimization, expanded multimodal model coverage, and a unified quantization framework.
Mar 28, 2026
- v2.0.2v2.0.2 — Smarter Listening, Better Testing
This release focuses on improving voice listening capabilities, testing reliability, and infrastructure upgrades.
Mar 17, 2026
- v1.37.1
This is a patch release that fixes a missing version bump for @logto/core-kit, which was causing downstream packages to reference non-existent exports.
Feb 28, 2026
- v0.16.0
v0.16.0 is a major alignment and capability release rebasing onto upstream vLLM v0.16.0. It features Qwen3-Omni/Qwen3-TTS performance and correctness improvements, MiMo-Audio production support, Bagel acceleration and scalability, Diffusion distributed execution expansion, and Quantization for DiT.
Feb 28, 2026
- v1.0.1
Feb 14, 2026
- v1.0.0
This is the initial production release marking the seed of the production branch for Cloudflare Pages deployment.
Feb 14, 2026
- v0.14.0
v0.14.0 is a feature-heavy release expanding diffusion/image-video generation and audio/TTS stacks. It features async chunk, stage-based deployment for Bagel, Qwen3-TTS support, Diffusion LoRA Adapter Support, DiT layerwise CPU offloading, and hardware platforms.
Jan 31, 2026
- v0.6.10
This release adds support for advanced AI models like OpenAI o1 and Gemini 2.0, enhances Redis capabilities with Sentinel and Cluster support, and introduces new providers such as DeepSeek and Replicate.
Feb 2, 2025
- v0.6.9
The release focuses on expanding provider support with SiliconFlow, Groq, and xAI, while improving Ollama capabilities and adding critical features like OIDC authentication and multipart request support.
Dec 22, 2024
- v0.6.8
This release introduces support for Spark4.0 Ultra and advanced features like Claude 3 function calling, while improving infrastructure with Cloudflare support and a new Proxy channel type for flexible routing.
Aug 6, 2024
- v0.6.7
Highlights include the addition of GPT-4o support, integration with ByteDance's Doubao, and the Tencent V3 API, alongside improvements to token generation and theme stability.
Jun 25, 2024