llama.cpp Releases
684 releases of ggml-org/llama.cpp
- b9604
Restores support for the SYCL backend by fixing CI build and release processes. Fixes include ccache usage, GitHub cache removal, and Windows build processor count settings.
Jun 12, 2026
- b9603
Added support for Q5 quantization types on Adreno GPUs.
Jun 12, 2026
- b9601
Build fix for Vulkan to address compilation issues with eMesaHoneykrisp.
Jun 11, 2026
- v3.7.0
Major release of PP-OCRv6 featuring significant accuracy improvements and multi-language support.
Jun 11, 2026
- b9596
Server log optimization for router mode and maintenance of various binary builds.
Jun 11, 2026
- b9594
Refactoring of vocabulary normalizer flags to improve code structure and add new options.
Jun 11, 2026
- b9592
Security update updating the vendor library LibreSSL to version 4.3.2.
Jun 10, 2026
- b9591
Optimization of MTP (Multi-Token Prediction) recurrent state handling to reduce memory overhead.
Jun 10, 2026
- b9590
Bug fix for LFM2/LFM2.5 template handler to ensure JSON schema is respected.
Jun 10, 2026
- v0.6.2
Release of WeKnora v0.6.2 featuring HNSW index updates and frontend enhancements.
Jun 10, 2026
- v0.5.6
Release of microsandbox v0.5.6 with improvements to CLI, runtime stability, and metrics.
Jun 10, 2026
- v2.0.3
This release introduces a comprehensive video creation pipeline featuring high-quality dubbing, customizable subtitle styles, and cover generation capabilities.
Jun 9, 2026
- v2.0.0SANA-Video && SANA-WM
This major release introduces SANA-Video and SANA-WM (Watermark) features, alongside significant updates for training, inference, and integration with external tools like ComfyUI.
Jun 9, 2026
- b9553
This release improves sampler name recognition by removing a flag that restricted matching. The system now always matches canonical and alternative names in a case-insensitive manner, fixing issues with the llama-server UI.
Jun 7, 2026
- b9551
This release optimizes the key-value (KV) cache memory management by avoiding unnecessary cell copies, aiming to improve efficiency and reduce overhead.
Jun 7, 2026
- b9550
This release fixes a crash/bug where oversized assistant views caused tensor overflow during graph reservation by ensuring the shared KV cache size follows the source cache size.
Jun 7, 2026
- b9549
This release adds support for Gemma4 Multi-Token Prediction (MTP) capability to enable more efficient text generation.
Jun 7, 2026
- b9548
This release fixes a vocabulary compatibility check issue that likely ensures models load correctly across different versions or configurations.
Jun 7, 2026
- v0.5.5
This release focuses on improving sandbox management through labels and metrics, introduces SSH capabilities, and adds Python support. It includes breaking changes to the Go SDK's sandbox creation API and adds Homebrew support for macOS users.
Jun 5, 2026
- v0.5.4
Updates the microsandbox runtime with boot argument limits, idle timeout fixes, and archive support.
Jun 1, 2026