v0.31.2
Notifuse/notifusev0.31.2Jul 6, 2026by github-actions[bot]
AI Summary
Enabled flash attention on older NVIDIA GPUs and added iGPU offloading for vision models, while fixing several bugs related to structured outputs and model loading.
Key Highlights
- Enabled flash attention on older NVIDIA GPUs (compute capability 6.x)
- iGPU can now offload vision models with padding to fit available memory
- Fixed structured output for thinking models when thinking is disabled
- Disabled telemetry for Claude Code by default
New Features
- iGPU offloading for vision models with padding to fit available memory
Full Release Notes
## What's Changed * Enabled flash attention on older NVIDIA GPUs (compute capability 6.x) * iGPU can now offload vision models with padding to fit available memory * Fixed structured output for thinking models when thinking is disabled * Hardened GGUF model creation * `ollama launch` for Claude Code now disables telemetry by default * Fixed loading models on paths with non-UTF-8 characters * Updated the MLX and llama.cpp engines ## New Contributors * @kevinpark1217 made their first contribution in https://github.com/ollama/ollama/pull/16949 **Full Changelog**: https://github.com/ollama/ollama/compare/v0.31.1...v0.31.2