v4.6.1
MontFerret/ferretv4.6.1Jul 6, 2026by mudler
AI Summary
This release focuses on stability improvements and new API capabilities, specifically adding a models capabilities endpoint and logging OAuth errors. It also introduces Prometheus metrics for agents and fixes VRAM constraints to prevent out-of-memory issues.
Key Highlights
- Added GET /v1/models/capabilities endpoint
- Enhanced auth logging to show the real cause of OIDC/OAuth failures
- Added native Prometheus metrics for agent chat runs
- Fixed VRAM constraints by capping auto-derived context and embedding batches
- Updated core dependencies (llama.cpp, ASR libraries)
New Features
- GET /v1/models/capabilities endpoint
- Prometheus metrics for agent chat runs
- Improved auth logging
- VRAM constraint fixes
Full Release Notes
<!-- Release notes generated using configuration in .github/release.yml at master --> ## What's Changed ### Other Changes * fix(auth): log the real cause of OIDC/OAuth user-info failures by @localai-bot in https://github.com/mudler/LocalAI/pull/10679 * feat(api): add GET /v1/models/capabilities endpoint by @localai-bot in https://github.com/mudler/LocalAI/pull/10687 * chore: :arrow_up: Update CrispStrobe/CrispASR to `1109cb3fcae2e242c2b3d42ec0e3fd6e813f2ce7` by @localai-bot in https://github.com/mudler/LocalAI/pull/10685 * chore: :arrow_up: Update ggml-org/llama.cpp to `665892536dfb1b7532161e3182304bd35c33e768` by @localai-bot in https://github.com/mudler/LocalAI/pull/10681 * chore(model-gallery): :arrow_up: update checksum by @localai-bot in https://github.com/mudler/LocalAI/pull/10686 * chore: :arrow_up: Update vllm-metal (darwin) to `v0.3.0.dev20260704102955` by @localai-bot in https://github.com/mudler/LocalAI/pull/10668 * fix(ui): center the home empty-state wizard by @localai-bot in https://github.com/mudler/LocalAI/pull/10691 * docs: :arrow_up: update docs version mudler/LocalAI by @localai-bot in https://github.com/mudler/LocalAI/pull/10680 * chore: :arrow_up: Update CrispStrobe/CrispASR to `09df654e304947f7521e1f52992ceacccf03c300` by @localai-bot in https://github.com/mudler/LocalAI/pull/10693 * chore: :arrow_up: Update ServeurpersoCom/omnivoice.cpp to `daedb763fd442e0916eb130a479fdd74947291c0` by @localai-bot in https://github.com/mudler/LocalAI/pull/10682 * chore: :arrow_up: Update ServeurpersoCom/qwentts.cpp to `73fe0c67bbf0898ba2999535e0680a02a7f8537d` by @localai-bot in https://github.com/mudler/LocalAI/pull/10683 * feat(agents): native Prometheus metrics for agent chat runs by @walcz-de in https://github.com/mudler/LocalAI/pull/10689 * chore: :arrow_up: Update ggml-org/llama.cpp to `2da668617612d2df773f966e3b0ee22dc2beef7b` by @localai-bot in https://github.com/mudler/LocalAI/pull/10694 * fix(reasoning): don't persist request-scoped reasoning_effort as an operator disable (#10622) by @Anai-Guo in https://github.com/mudler/LocalAI/pull/10623 * fix(startup): scope generated-content and upload dirs to the current user by @localai-bot in https://github.com/mudler/LocalAI/pull/10698 * fix(config): cap auto-derived context to fit VRAM by @localai-bot in https://github.com/mudler/LocalAI/pull/10696 * fix(llama-cpp): cap single-pass embedding batch to fit VRAM by @localai-bot in https://github.com/mudler/LocalAI/pull/10695 **Full Changelog**: https://github.com/mudler/LocalAI/compare/v4.6.0...v4.6.1