v0.9.0

mantinedev/mantinev0.9.0Jul 7, 2026by EricLBuehler

AI Summary

This release focuses on significant performance optimizations across Metal, CUDA, and CPU backends, alongside the addition of support for LFM 2.5 models (Dense, MoE, and VL). It also introduces improvements to the distributed training system, multi-model planning, and the installation workflow.

Key Highlights

  • Support for LFM 2.5 models including Dense, MoE, and VL variants.
  • Significant performance improvements for Metal runtime and AOT compilation.
  • CPU optimization for both x86 and aarch64 architectures.
  • Enhanced distributed training with sophisticated TP fallback and resolution.
  • Multi-model support planning capabilities.

New Features

  • Metal runtime and AOT compilation system improvements
  • LFM 2.5 model support (dense, MoE, and VL)
  • Multi-model support planning
  • CUDA PTX module preloading
  • Aarch64 CPU repacking kernels and performance improvements

Full Release Notes

## What's Changed
* Fix CUDA version extraction in install.sh by @zaqxsw-dev in https://github.com/EricLBuehler/mistral.rs/pull/2287
* feat(metal): improve runtime and aot compilation system by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2288
* feat(release): split cuda workers for docker release by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2289
* feat(install): cleanup install system for updating by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2290
* feat(distributed): more sophisticated TP fallback and resolution system by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2291
* chore(mistralrs-core): cap tojson indent and remove chat template panic paths by @eharris128 in https://github.com/EricLBuehler/mistral.rs/pull/2286
* feat(quant): improve uqff resolution per model ids by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2293
* Fix agentic loop not starting for registered tool callbacks by @aeiklorvy in https://github.com/EricLBuehler/mistral.rs/pull/2279
* feat(models): Support LFM 2.5 (dense + MoE) and LFM 2.5 VL! by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2295
* feat(cuda): preload cuda ptx modules by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2296
* refactor(core): move code exec shims from lib.rs to proper location by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2297
* docs(models): improve supported models readability by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2298
* feat(multi-model): support planning for loading multi-model by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2299
* feat(cuda): update cutlass git revision by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2300
* feat(cpu): aarch64 cpu repacking kernels and performance improvements by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2304
* feat(cpu): optimize decode and prefill performance on x86 and aarch64 by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2311
* fix(tune): detect CUDA backend on Windows by @matt-minev in https://github.com/EricLBuehler/mistral.rs/pull/2312
* fix(server): keep Responses output item IDs stable by @matt-minev in https://github.com/EricLBuehler/mistral.rs/pull/2309

## New Contributors
* @zaqxsw-dev made their first contribution in https://github.com/EricLBuehler/mistral.rs/pull/2287
* @eharris128 made their first contribution in https://github.com/EricLBuehler/mistral.rs/pull/2286
* @aeiklorvy made their first contribution in https://github.com/EricLBuehler/mistral.rs/pull/2279
* @matt-minev made their first contribution in https://github.com/EricLBuehler/mistral.rs/pull/2312

**Full Changelog**: https://github.com/EricLBuehler/mistral.rs/compare/v0.8.23...v0.9.0