inference Releases
114 releases of xorbitsai/inference
- v0.14.4
Adds support for CogVideoX-5b, enhances padding for SD inpainting models, moves Matcha to third-party, and fixes various bugs related to model launching, health checks, and configuration keys.
Aug 30, 2024
- v0.14.3
Introduces ChatTTS voice support, SD3-medium inpainting, Fish Speech model, and CogVLM2-video. Adds UI parameters for models and LMDeploy support for InternVL2 video capabilities.
Aug 25, 2024
- v0.14.2
Adds support for Gemma-2-it and InternLM2.5, enables FP8 support for vLLM and SGLang, and enhances InternVL2 video support. Removes some old builtin models.
Aug 16, 2024
- v0.14.1.post1
A minor post-release patch fixing AutoAWQ version limits to resolve Docker issues and updating documentation.
Aug 13, 2024
- v0.14.1
Adds support for SenseVoice audio-to-text, Flux.1 models, Kolors image model, CogVideoX video model, and MiniCPM-v-2_6. Enhances internal server error handling and adds streaming option to Benchmark.
Aug 9, 2024
- v0.14.0.post1
Post-release hotfixes addressing internal server errors, flexible model registration issues, and UI bugs related to model_path configuration.
Aug 5, 2024
- v0.14.0
Major update adding support for launching models via model_path, gte-Qwen2-7B-instruct multi-GPU deployment, SGLang support for Llama 3 and Qwen 2, and image-to-image capabilities.
Aug 2, 2024
- v0.13.3
Adds GLM4 stream tool calls, support for csg-wukong-chat, mistral-nemo-instruct, mistral-large-instruct, CosyVoice speech, llama-3.1, and rembg for background removal.
Jul 26, 2024
- v0.13.2
This release introduces support for SD inpainting, CodeGeeX4, and InternLM2.5-chat models, along with streaming capabilities for ChatTTS. It also includes bug fixes for Chinese character encoding and stream errors in vLLM and sglang backends.
Jul 19, 2024
- v0.13.1
This update adds support for Flexible Models and enhances the UI by allowing users to specify download hubs. It also updates ChatTTS and removes the `chatglm-cpp` dependency to fix issues with `llama-cpp-python`.
Jul 12, 2024
- v0.13.0
A major update that introduces the MLX engine for Apple Silicon, GGUF file support for Qwen2, and continuous batching for vision models. It also adds Aliyun Docker image support and various stability improvements.
Jul 5, 2024
- v0.12.3
This release adds SD3 support, Jina Rerank v2, and Tensorizer integration. It also introduces UI features like favorites and automatic configuration retrieval, along with cluster deletion functionality.
Jun 28, 2024
- v0.12.2.post1
A hotfix version that pins the `chatglm-cpp` version to `v0.3.x` to ensure stability and compatibility.
Jun 22, 2024
- v0.12.2
This update focuses on tool support for Qwen MOE models and continuous batching for all models using the transformers backend. It also improves UI capabilities for custom models and rerank token usage.
Jun 21, 2024
- v0.12.1
Adds tool call support for Qwen2 and GLM4-chat, along with a new method to download models from CSGHub. The UI also gains viewing and deleting cache data capabilities, and GLM-4V gains quantization support.
Jun 14, 2024
- v0.12.0
A major feature release introducing several new models including Mini-CPM-Llama3, GLM4, Mistral, Codestral, and Qwen2. It also adds ChatTTS support and continuous batching for chat models on the transformers backend.
Jun 7, 2024
- v0.11.3
Introduces support for Yi-1.5-chat-16k, CogVLM, and Telechat models. Enhancements include adding engine options to launch details and a real paths column in the UI.
May 31, 2024
- v0.11.2.post1
A hotfix version specifically addressing and fixing the launch model error encountered when using torch 2.3.0.
May 24, 2024
- v0.11.2
This release adds support for new models like DeepSeek LLM and CodeQwen1.5, along with a new memory calculation command. It also enhances functionality for querying cached models and updates compatibility with various libraries.
May 24, 2024
- v0.11.1
This release focuses on stability by adding support for the Yi-1.5 series and refining LoRa adaptation methods. It includes fixes for Docker images, `llama.cpp` models, and various UI enhancements.
May 17, 2024