v0.30.8

sampotts/plyrv0.30.8Jun 12, 2026by github-actions[bot]

AI Summary

This release focuses on enhancing MLX (Apple Silicon) inference stability and performance, alongside improvements to prompt caching and model support.

Key Highlights

  • Improved prompt caching by decoupling it from context shift for better KV cache reuse
  • More stable MLX inference with hardened linear and embedding layers
  • MLX runner creates snapshots during prompt processing for improved reliability
  • Fixed `ollama launch` selecting the wrong provider in some cases
  • Improved recurrent model support with per-boundary states from gated-delta kernels

New Features

  • Improved recurrent model support
  • Enhanced prompt caching mechanism

Full Release Notes

## What's Changed
* Fixed `ollama launch` selecting the wrong provider in some cases
* Improved prompt caching by decoupling it from context shift for better KV cache reuse
* More stable MLX inference with hardened linear and embedding layers
* MLX runner now creates snapshots during prompt processing and speculative decoding for improved reliability
* Improved recurrent model support with per-boundary states from the gated-delta kernels

**Full Changelog**: https://github.com/ollama/ollama/compare/v0.30.7...v0.30.8