mistral.rs Releases
113 releases of EricLBuehler/mistral.rs
- v1.2.01.2.0
PentAGI 1.2.0 delivers major AI capabilities including support for latest reasoning models (Gemini 2.5/3.0, Claude Sonnet 4+, DeepSeek R1, OpenAI o-series), token caching for 40-70% cost reduction, REST API access with JWT authentication, and comprehensive Langfuse v3 observability.
Feb 25, 2026
- v0.7.0
Major release introducing the new mistralrs-cli, Prefix Caching for PagedAttention, and significant model expansion including Embedding Gemma, Qwen 3 Embedding, Gemma 3n, GLM-4, and Granite Hybrid MoE. Added dynamic model loading for runtime model management, CUDA 13.0/13.1 support with highly optimized fused kernels, and migrated to stable candle 0.9.2.
Jan 28, 2026
- v1.1.01.1.0
PentAGI 1.1.0 focuses on bug fixes and improvements including LiteLLM passthrough support for all providers, Windows file path compatibility through container path migration, Ollama single model configuration, and installer v1.0.0 with comprehensive stability improvements.
Jan 17, 2026
- 1.4.5
This release introduces new remote control permission settings, a relative mouse mode, and mobile-specific terminal improvements, alongside various bug fixes for window positioning and file transfers.
Jan 9, 2026
- v1.0.11.0.1
This update focuses on improving AI agent reliability through enhanced error diagnostics, a stable DuckDuckGo search implementation, and updated OpenAI provider configurations to handle content filters and instability.
Jan 6, 2026
- v0.15.0
This release focuses on architectural refactoring of workers and queues, enhancing user experience with inline error messages across multiple forms, and expanding functionality with flow import/export capabilities and new app integrations like Anthropic and Signalwire. It also includes significant improvements to test coverage and database schema management.
Aug 8, 2025
- v0.6.0
Feature-rich release with Llama 4 support, Qwen 3/MoE/VL models, multimodal prefix caching, and a new web chat application. Added Fast sampler, CPU FlashAttention, MCP support, and expanded vision/audio capabilities including SIGLIP, Dia 1.6b TTS, and Phi-4MM audio.
Jun 10, 2025
- v0.5.0
Major release adding support for Gemma 3, Qwen 2.5 VL, Mistral Small 3.1, and Phi 4 Multimodal. Introduced native tool calling for multiple model families, Tensor Parallelism via NCCL, FlashAttention V3 integration, and revamped prefix cacher system.
Mar 24, 2025
- v1.1.2
A minor update focusing on fixing schema validation issues and updating the Google AI SDK dependency.
Feb 21, 2025
- v1.1.1
A patch release specifically addressing the schema validation issue in nested fields.
Feb 20, 2025
- v0.4.0
Release adding DeepSeek V2, V3, R1, and MiniCpm-O 2.6 models. Introduced imatrix quantization, automatic device mapping, BNB quantization, Metal PagedAttention, and blockwise FP8 dequantization.
Jan 22, 2025
- v1.0.12
Adds support for multiple file formats and expands schema capabilities to include boolean and enum types.
Jan 14, 2025
- v1.0.11
Adds auto-generation of schemas and introduces new file converters and formatters.
Dec 14, 2024
- v0.3.4
Release adding Qwen2-VL and Idefics 3/SmolVLM support with a 6x prompt performance boost. Introduced more efficient non-PagedAttention KV cache and public tokenization API.
Nov 28, 2024
- v0.14.0
This release introduces new integrations for MailerLite, Mailchimp, Jotform, and ClickUp, alongside a comprehensive suite of REST API endpoints for managing flows, users, and connections. It also includes significant UI enhancements and bug fixes.
Nov 18, 2024
- v0.3.2
A feature release focusing on quantization improvements with FP8 ISQ and GPTQ Marlin support, a significant performance boost on Metal, and the addition of Python package wheels along with support for Qwen 2.5.
Oct 28, 2024
- v0.3.1
Introduces UQFF format, FLUX diffusion model, and Llama 3.2 Vision support, with MSRV bumped to 1.79.0.
Sep 29, 2024
- v0.3.0
Introduces ISQ topology and device mapping, enhances FlashAttention performance, and removes plotly dependencies. MSRV 1.79.0.
Sep 2, 2024
- v0.2.5
Improves ISQ parsing, implements prompt chunking, and adds GPTQ and HQQ quantization support.
Aug 16, 2024
- v0.13.1
A minor patch release focused on dependency updates and specific bug fixes, including the removal of PWA support and the exposure of installation completion status.
Aug 2, 2024