v0.9.0

stackblitz/bolt.newv0.9.0Jul 7, 2026by EricLBuehler

AI Summary

A performance and feature-focused release introducing support for the LFM 2.5 model family, optimizing CPU and GPU backends, and refining the distributed training system with sophisticated tensor parallelism (TP) fallback.

Key Highlights

  • Support for LFM 2.5 (dense + MoE) and LFM 2.5 VL
  • CPU Performance Optimizations (x86 and aarch64)
  • Distributed Training Improvements (TP fallback)
  • Metal Runtime Enhancements (macOS)
  • CUDA Preloading and Version Fixes

New Features

  • Support for LFM 2.5 (dense + MoE) and LFM 2.5 VL
  • Multi-model support for planning
  • Distributed training with more sophisticated TP fallback
  • Metal runtime and AOT compilation improvements
  • CUDA version extraction and preload fixes
  • CPU repacking kernels and performance improvements

Full Release Notes

## What's Changed
* Fix CUDA version extraction in install.sh by @zaqxsw-dev in https://github.com/EricLBuehler/mistral.rs/pull/2287
* feat(metal): improve runtime and aot compilation system by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2288
* feat(release): split cuda workers for docker release by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2289
* feat(install): cleanup install system for updating by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2290
* feat(distributed): more sophisticated TP fallback and resolution system by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2291
* chore(mistralrs-core): cap tojson indent and remove chat template panic paths by @eharris128 in https://github.com/EricLBuehler/mistral.rs/pull/2286
* feat(quant): improve uqff resolution per model ids by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2293
* Fix agentic loop not starting for registered tool callbacks by @aeiklorvy in https://github.com/EricLBuehler/mistral.rs/pull/2279
* feat(models): Support LFM 2.5 (dense + MoE) and LFM 2.5 VL! by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2295
* feat(cuda): preload cuda ptx modules by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2296
* refactor(core): move code exec shims from lib.rs to proper location by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2297
* docs(models): improve supported models readability by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2298
* feat(multi-model): support planning for loading multi-model by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2299
* feat(cuda): update cutlass git revision by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2300
* feat(cpu): aarch64 cpu repacking kernels and performance improvements by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2304
* feat(cpu): optimize decode and prefill performance on x86 and aarch64 by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2311
* fix(tune): detect CUDA backend on Windows by @matt-minev in https://github.com/EricLBuehler/mistral.rs/pull/2312
* fix(server): keep Responses output item IDs stable by @matt-minev in https://github.com/EricLBuehler/mistral.rs/pull/2309

## New Contributors
* @zaqxsw-dev made their first contribution in https://github.com/EricLBuehler/mistral.rs/pull/2287
* @eharris128 made their first contribution in https://github.com/EricLBuehler/mistral.rs/pull/2286
* @aeiklorvy made their first contribution in https://github.com/EricLBuehler/mistral.rs/pull/2279
* @matt-minev made their first contribution in https://github.com/EricLBuehler/mistral.rs/pull/2312

**Full Changelog**: https://github.com/EricLBuehler/mistral.rs/compare/v0.8.23...v0.9.0