v0.9.0
ggml-org/llama.cppv0.9.0Jul 7, 2026by EricLBuehler
AI Summary
Note: This release appears to be from mistral.rs repository, not ggml-org/llama.cpp. It includes multiple features including Metal runtime improvements, CUDA worker splitting, distributed training enhancements, LFM 2.5 model support, and extensive performance optimizations.
Key Highlights
- Metal runtime and AOT compilation system improvements
- CUDA worker splitting for Docker releases
- LFM 2.5 model support (dense + MoE and VL variants)
- Distributed training fallback and resolution system
- CPU optimizations for both x86 and aarch64
New Features
- LFM 2.5 model support
- Metal runtime improvements
- CUDA worker splitting
- Distributed training enhancements
- CPU performance optimizations
Full Release Notes
## What's Changed * Fix CUDA version extraction in install.sh by @zaqxsw-dev in https://github.com/EricLBuehler/mistral.rs/pull/2287 * feat(metal): improve runtime and aot compilation system by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2288 * feat(release): split cuda workers for docker release by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2289 * feat(install): cleanup install system for updating by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2290 * feat(distributed): more sophisticated TP fallback and resolution system by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2291 * chore(mistralrs-core): cap tojson indent and remove chat template panic paths by @eharris128 in https://github.com/EricLBuehler/mistral.rs/pull/2286 * feat(quant): improve uqff resolution per model ids by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2293 * Fix agentic loop not starting for registered tool callbacks by @aeiklorvy in https://github.com/EricLBuehler/mistral.rs/pull/2279 * feat(models): Support LFM 2.5 (dense + MoE) and LFM 2.5 VL! by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2295 * feat(cuda): preload cuda ptx modules by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2296 * refactor(core): move code exec shims from lib.rs to proper location by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2297 * docs(models): improve supported models readability by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2298 * feat(multi-model): support planning for loading multi-model by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2299 * feat(cuda): update cutlass git revision by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2300 * feat(cpu): aarch64 cpu repacking kernels and performance improvements by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2304 * feat(cpu): optimize decode and prefill performance on x86 and aarch64 by @EricLBuehler in https://github.com/EricLBuehler/mistral.rs/pull/2311 * fix(tune): detect CUDA backend on Windows by @matt-minev in https://github.com/EricLBuehler/mistral.rs/pull/2312 * fix(server): keep Responses output item IDs stable by @matt-minev in https://github.com/EricLBuehler/mistral.rs/pull/2309 ## New Contributors * @zaqxsw-dev made their first contribution in https://github.com/EricLBuehler/mistral.rs/pull/2287 * @eharris128 made their first contribution in https://github.com/EricLBuehler/mistral.rs/pull/2286 * @aeiklorvy made their first contribution in https://github.com/EricLBuehler/mistral.rs/pull/2279 * @matt-minev made their first contribution in https://github.com/EricLBuehler/mistral.rs/pull/2312 **Full Changelog**: https://github.com/EricLBuehler/mistral.rs/compare/v0.8.23...v0.9.0