inference Releases
114 releases of xorbitsai/inference
- v2.3.0
This release adds Qwen3.5 support for vLLM and transformers, introduces seed and repetition_penalty parameters, removes qwen2-audio model, and fixes various bugs related to worker initialization and vLLM.
Mar 13, 2026
- v0.7.382
This release focuses on enhancing process automation and document management capabilities. It introduces specific features for field locking and the ability to revert document updates within workflows. Additionally, the update includes various stability improvements and bug fixes to refine the user experience.
Mar 2, 2026
- v2.2.0
This release adds support for DeepSeek V3.1 tool parser, introduces new model support including Kimi-K2.5, MiniMax-M2.5, and glm-5, and fixes various bugs.
Feb 28, 2026
- v0.7.375
This release significantly expands localization options and workflow management features. Users can now import and export processes, and guests have enhanced collaboration capabilities. The update also features infrastructure improvements and a redesigned workspace join experience.
Feb 23, 2026
- v2.1.0
Release v2.1.0 adds support for multiple new models including GLM-4.7, FLUX.2-klein, Qwen3-ASR, and MinerU2.5. Includes bug fixes for vllm embedding/reranker models and refactoring of API schemas into domain-specific modules.
Feb 14, 2026
- v2.0.0
Major version 2.0.0 release introduces Qwen3-VL family models, MinerU 2.5 OCR model, GLM-4.6, and Z-Image. Adds video GGUF cache support, chat_template.jinja support, and UI improvements with backend API data-driven approach.
Jan 31, 2026
- v1.17.1
Hotfix version addressing issues from v1.17.0. No specific changes detailed in release notes.
Jan 13, 2026
- v1.17.0
Version 1.17.0 adds enable_thinking argument support, MThreads (MUSA) GPU support, minimax tool call support, and new Qwen image models. Includes FP4 quantization support, video GGUF support, and multi-engine configurations.
Jan 10, 2026
- v1.16.0
Version 1.16.0 adds DeepSeek-V3.2-Exp support, Fun-ASR speech recognition models, Qwen-Image-Layered, and MiniMaxM2. Introduces vacc support, continuous batching for MLX chat models, and rerank async batch functionality.
Dec 27, 2025
- v1.15.0
This release adds support for DeepSeek-V3.2, PaddleOCR-VL, Z-Image-Turbo models, introduces multi-replica on single GPU with launch strategy, and enhances tool calls support.
Dec 13, 2025
- v1.14.0
This release adds vLLM 0.11.1+ compatibility, introduces new virtualenv APIs, adds support for rerank model for llamacpp, and enhances startup parallelization.
Nov 30, 2025
- v3.20.23.20.2
This is a maintenance release focused on fixing a specific bug where a NoneType object caused an error when attempting to access the lower() method.
Nov 27, 2025
- v3.20.13.20.1
Nov 22, 2025
- v3.20.03.20.0
This major release introduces significant enhancements to the YouTube downloading functionality, including new runtime support, custom configuration options, and UI improvements.
Nov 22, 2025
- v1.13.0
This release adds Qwen3-VL-MLX support, introduces auto batch embedding, updates models via Xinference model hub, and enhances IndexTTS2 capabilities.
Nov 15, 2025
- v1.12.0
This release adds support for jina-reranker-v3, qwen3-omni, and DeepSeek-OCR models, introduces OCR gradio UI, and adds Python 3.13 support.
Nov 2, 2025
- v1.11.0.post1
This is a post-release patch for v1.11.0 that addresses specific bugs in Qwen3 model transformers and UI progress bar display issues.
Oct 20, 2025
- v1.11.0
This release introduces comprehensive support for Qwen3 models, OpenAI image edit API, and vLLM multi-model capabilities, along with various enhancements and bug fixes.
Oct 19, 2025
- v1.10.1
This release focuses on OpenAI API enhancements, UI improvements, and support for several new models including IndexTTS2, Qwen3-VL, and Qwen3-Next.
Oct 1, 2025
- v1.10.0
This release introduces IP restriction functionality, Anthropic API format support, and enhanced JSON schema output for vLLM, along with various bug fixes.
Sep 13, 2025