inference Releases
114 releases of xorbitsai/inference
- v0.11.0
A major update requiring users to specify a `model_engine` when launching models. It introduces support for Mixtral-8x22b, Phi-3, and Ascend hardware, along with significant performance improvements and refactoring.
May 11, 2024
- v0.10.3
This release adds support for the Llama-3 family and the Belle-whisper-large-v3-zh audio model. It includes bug fixes for model launching and cache management.
Apr 24, 2024
- v0.10.2.post1
A minor patch release fixing dependency issues in the client and RESTful client packages to ensure proper functioning.
Apr 19, 2024
- v0.10.2
This release introduces replica configuration for embeddings and rerank models, supports multi-LoRA usage, and adds several new models like SeaLLM-7B and BGE rerankers.
Apr 19, 2024
- v0.10.1
This release adds support for Qwen1.5 32B chat and Qwen MoE models, with improvements to streaming tool calls and multi-GPU support.
Apr 12, 2024
- v0.10.0
A significant update introducing the Web UI for audio models, OmniLMM support, and vLLM support for DeepSeek models. It also introduces an OAuth system.
Mar 29, 2024
- v0.9.4
This release adds support for the CodeShell model and introduces SGLang backend support. It also updates vLLM to support the latest models.
Mar 21, 2024
- v0.9.3
This release adds the Yi-9B model and image generation functionality. It improves OpenAI API compatibility for the model list and removes quantization limits for Apple Metal.
Mar 15, 2024
- v0.9.2
This release focuses on distributed inference capabilities and model flexibility. It introduces new features for GGUF file handling and LoRA support, alongside UI improvements for GPU layer management.
Mar 8, 2024
- v0.9.1
v0.9.1 adds CPU-only Docker support and improves model downloading from ModelScope. It also refines the RESTful client and command-line interface for better usability.
Mar 1, 2024
- v0.9.0
The v0.9.0 major release introduces Intel GPU support and refactors device-related code. It also adds the Gemma model series and improves the UI to show cluster resources.
Feb 22, 2024
- v0.8.5
This release features the implementation of a web UI for text-to-image models and adds support for the Qwen 1.5 series. It also upgrades to Pydantic v2.
Feb 6, 2024
- v0.8.4
v0.8.4 focuses on bug fixes and documentation updates. It includes UI improvements for long model names and adds GGUF models for llama-2-chat.
Feb 4, 2024
- 2.4.02.4.0 - Gone Phishing
This release focuses on enhancing campaign management and security, introducing features like HTML templates, custom hostnames, and SOCKS5/HTTP(S) proxy support. It also improves parameter handling by enforcing URL encoding and adds import/export capabilities for URLs.
Sep 14, 2020