v0.30.8
ollama/ollamav0.30.8Jun 12, 2026by github-actions[bot]
AI Summary
Enhances MLX inference stability with snapshots, improves prompt caching by decoupling it from context shifts, and fixes provider selection issues in `ollama launch`.
Key Highlights
- Improved prompt caching with better KV cache reuse
- MLX inference reliability improvements
- Fixed `ollama launch` provider selection
New Features
- Decoupled prompt caching from context shift
- MLX snapshots during processing
- Recurrent model support improvements
Full Release Notes
## What's Changed * Fixed `ollama launch` selecting the wrong provider in some cases * Improved prompt caching by decoupling it from context shift for better KV cache reuse * More stable MLX inference with hardened linear and embedding layers * MLX runner now creates snapshots during prompt processing and speculative decoding for improved reliability * Improved recurrent model support with per-boundary states from the gated-delta kernels **Full Changelog**: https://github.com/ollama/ollama/compare/v0.30.7...v0.30.8