v0.30.8
sampotts/plyrv0.30.8Jun 12, 2026by github-actions[bot]
AI Summary
This release focuses on enhancing MLX (Apple Silicon) inference stability and performance, alongside improvements to prompt caching and model support.
Key Highlights
- Improved prompt caching by decoupling it from context shift for better KV cache reuse
- More stable MLX inference with hardened linear and embedding layers
- MLX runner creates snapshots during prompt processing for improved reliability
- Fixed `ollama launch` selecting the wrong provider in some cases
- Improved recurrent model support with per-boundary states from gated-delta kernels
New Features
- Improved recurrent model support
- Enhanced prompt caching mechanism
Full Release Notes
## What's Changed * Fixed `ollama launch` selecting the wrong provider in some cases * Improved prompt caching by decoupling it from context shift for better KV cache reuse * More stable MLX inference with hardened linear and embedding layers * MLX runner now creates snapshots during prompt processing and speculative decoding for improved reliability * Improved recurrent model support with per-boundary states from the gated-delta kernels **Full Changelog**: https://github.com/ollama/ollama/compare/v0.30.7...v0.30.8