v0.31.2

Notifuse/notifusev0.31.2Jul 6, 2026by github-actions[bot]

AI Summary

Enabled flash attention on older NVIDIA GPUs and added iGPU offloading for vision models, while fixing several bugs related to structured outputs and model loading.

Key Highlights

  • Enabled flash attention on older NVIDIA GPUs (compute capability 6.x)
  • iGPU can now offload vision models with padding to fit available memory
  • Fixed structured output for thinking models when thinking is disabled
  • Disabled telemetry for Claude Code by default

New Features

  • iGPU offloading for vision models with padding to fit available memory

Full Release Notes

## What's Changed

* Enabled flash attention on older NVIDIA GPUs (compute capability 6.x)
* iGPU can now offload vision models with padding to fit available memory
* Fixed structured output for thinking models when thinking is disabled
* Hardened GGUF model creation
* `ollama launch` for Claude Code now disables telemetry by default
* Fixed loading models on paths with non-UTF-8 characters
* Updated the MLX and llama.cpp engines

## New Contributors
* @kevinpark1217 made their first contribution in https://github.com/ollama/ollama/pull/16949

**Full Changelog**: https://github.com/ollama/ollama/compare/v0.31.1...v0.31.2