v0.1.46-beta

oumi-ai/oumiv0.1.46-betaJun 12, 2026by shimmyshimmer

AI Summary

This release introduces significant enhancements to Unsloth Studio, including support for DiffusionGemma and Gemma 4 MTP with audio capabilities, a new model Hub for managing Hugging Face assets, and experimental RAG features for chatting with local files. It also improves hardware compatibility across platforms and optimizes local tool calling and API behaviors.

Key Highlights

  • DiffusionGemma and Gemma 4 MTP integration with audio support.
  • New Model Hub for browsing, downloading, and managing local/dataset assets.
  • Experimental Chat with Files (RAG) feature for local documents.
  • Enhanced llama.cpp support with Tensor Parallelism (+30% throughput).
  • Improved hardware compatibility (Strix Halo ROCm, Blackwell CUDA).

New Features

  • DiffusionGemma support
  • Gemma 4 MTP support
  • Audio chat for Gemma 4
  • Model Hub page
  • Chat with Files / RAG
  • Cloudflare HTTPS free tunnels
  • Tensor parallelism for GGUFs
  • Custom provider option in Connections

Full Release Notes

We've merged over 150 PRs this week so lots of new updates, a new model Hub and look!

### DiffusionGemma + Gemma 4 MTP + Audio
- Run and train [DiffusionGemma](https://unsloth.ai/docs/models/diffusiongemma) via [Unsloth Studio](https://unsloth.ai/docs/new/studio).
- [Gemma 4 MTP](https://unsloth.ai/docs/models/mtp) is here! Run [Gemma 4](https://unsloth.ai/docs/models/gemma-4) around 2x faster with MTP - MTP is auto enabled in Unsloth Studio.
- [Audio chat](https://unsloth.ai/docs/new/studio/chat) is now supported for Gemma 4 (`wav`, `mp3`, `m4a`,`flac`, `webm`).

### Hub + Download Manager (Experimental)
- Added a new **Hub** page for browsing, downloading, and managing Hugging Face models and datasets.
- Unsloth can now detect models and datasets already on your machine and show them alongside downloaded assets.
- Downloaded [GGUF models](https://unsloth.ai/docs/get-started/unsloth-model-catalog) now have direct **Run / New Chat** actions.

### Chat with Files / RAG (Experimental)
- Added **[Chat with Files](https://unsloth.ai/docs/new/studio/chat)** in Studio, letting you ask questions over your own documents and knowledge bases.
- Supports hybrid search, citations, PDF previews, per-thread documents, and a built-in `search_knowledge_base` tool.

### New Update Button + Hardware Support
- Unsloth now uses fresh, up-to-date [llama.cpp prebuilts](https://unsloth.ai/docs/new/changelog) across CUDA, ROCm, Windows, Linux, and macOS.
- Added an in-app **Update llama.cpp** button so users can update the local backend without reinstalling Studio.
- Improved Windows / WSL AMD support, [Strix Halo ROCm support](https://unsloth.ai/docs/get-started/install/amd), [Blackwell CUDA selection](https://unsloth.ai/docs/blog/fine-tuning-llms-with-blackwell-rtx-50-series-and-unsloth), and clearer installer messages.

### Local Chat, Tools & API Compatibility
- Local [tool calling](https://unsloth.ai/docs/basics/tool-calling-guide-for-local-llms) is more reliable, with better ordering of tool cards, fewer duplicate tool loops, and support for tool use with GGUF vision models.
- Improved [OpenAI-compatible API](https://unsloth.ai/docs/basics/inference-and-deployment/llama-server-and-openai-endpoint) and Anthropic-compatible API behavior for local Studio servers, including better errors, token usage, stop reasons, and [Claude Code compatibility](https://unsloth.ai/docs/basics/claude-code).

### Tool Calling, MCP, Encrypted Cloudflare Tunnels
- Bypass Permissions, Tool Call Permissions (Approve, Always Approve, Deny)
- 50% to 90% less tool call nudging issues without any accuracy loss
- MCP, Artifacts are now select-able
- Tensor parallelism is now enabled for GGUFs - get +30% throughput!
- Cloudflare HTTPS free tunnels is now added allowing for end to end encrypted studios!

### Training & General Fixes
- Improved [MLX support](https://unsloth.ai/docs/new/studio/install) with better model labels, generation speed stats, and fixes for [VLM training](https://unsloth.ai/docs/basics/vision-fine-tuning).
- Fixed several [training](https://unsloth.ai/docs/get-started/fine-tuning-llms-guide) and [dataset](https://unsloth.ai/docs/get-started/fine-tuning-llms-guide/datasets-guide) edge cases, including non-writable Hugging Face caches and custom dataset mappings.
- Added many UI polish fixes across chat, menus, model picker, dark mode, import/export, and settings.

### To update Unsloth or install a new Unsloth Studio, you must use:
**macOS, Linux, WSL:**
```
curl -fsSL https://unsloth.ai/install.sh | sh
```
**Windows:**
```
irm https://unsloth.ai/install.ps1 | iex
```
> [!WARNING]
> **DO NOT USE `unsloth studio update` since packaging will not get the latest updates**

## What's Changed
* Studio: llama.cpp update banner redesign, About tab license info, UI polish by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/6196
* Bump install.sh / install.ps1 pin to unsloth>=2026.6.3 by @danielhanchen in https://github.com/unslothai/unsloth/pull/6212
* Expose runtime context length for hub models by @alkinun in https://github.com/unslothai/unsloth/pull/6154
* Studio: fix llama.cpp update banner offering a downgrade / sticking on mix releases by @oobabooga in https://github.com/unslothai/unsloth/pull/6219
* Fix kwarg spacing in training files to satisfy pre-commit by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/6209
* Studio: reword the Cloudflare line when the public probe fails by @danielhanchen in https://github.com/unslothai/unsloth/pull/6217
* fix: deduplicate lemonade ROCm prebuilt selection log by @LeoBorcherding in https://github.com/unslothai/unsloth/pull/6021
* Stop false RoPE 'default' warning and fix rope drift gate on transformers 5 by @danielhanchen in https://github.com/unslothai/unsloth/pull/6223
* fix(studio): load run.py by path for editable installs by @jimdawdy-hub in https://github.com/unslothai/unsloth/pull/5909
* fix(studio): inherit llama_extra_args and honor --no-mmproj by @jimdawdy-hub in https://github.com/unslothai/unsloth/pull/5902
* fix(studio): adopt server-loaded model before chat auto-load by @jimdawdy-hub in https://github.com/unslothai/unsloth/pull/5900
* Fix stale sidebar regression test to match the gap-px markup by @danielhanchen in https://github.com/unslothai/unsloth/pull/6232
* Studio: gate the staged prebuilt runtime validation behind a flag (off by default) by @danielhanchen in https://github.com/unslothai/unsloth/pull/6216
* Fix FastModel config passthrough for sequence classification by @alkinun in https://github.com/unslothai/unsloth/pull/6203
* fix: decode subprocess output as UTF-8 in save.py on Windows by @dylanschroers in https://github.com/unslothai/unsloth/pull/6218
* patch: fix EmptyLogits gathering in nested payloads and Accelerate recursively_apply by @MdHussain121 in https://github.com/unslothai/unsloth/pull/6092
* Studio: show Apple GPU temperature and power in the GPU monitor (macOS) by @Ban921 in https://github.com/unslothai/unsloth/pull/6187
* Studio: Add inline confirmation (Allow/Always allow/Deny) for tool calls by @oobabooga in https://github.com/unslothai/unsloth/pull/5869
* Studio: guard Apple GPU power against negative counter-reset readings by @danielhanchen in https://github.com/unslothai/unsloth/pull/6235
* Fix step count mismatch when sequence packing is enabled by @IrakliXYZ in https://github.com/unslothai/unsloth/pull/5967
* fix/uv-bytecode-timeout by @alkinun in https://github.com/unslothai/unsloth/pull/6166
* Studio: tune llama.cpp env for data-center GPUs by @danielhanchen in https://github.com/unslothai/unsloth/pull/6098
* Studio: drop the on-disk freshness cache after a llama.cpp update by @danielhanchen in https://github.com/unslothai/unsloth/pull/6234
* Add missing RAG deps to no-torch Studio runtime requirements by @danielhanchen in https://github.com/unslothai/unsloth/pull/6236
* Studio: rounded rectangle hover states for menu items instead of pills by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/6210
* docs: repository cleanup by @Agnibha007 in https://github.com/unslothai/unsloth/pull/5617
* Run cross-platform parity test on Windows and macOS in CI by @danielhanchen in https://github.com/unslothai/unsloth/pull/6241
* chore(studio/frontend): normalize line endings to LF by @danielhanchen in https://github.com/unslothai/unsloth/pull/6012
* fix: respect absolute export paths to prevent cross-drive copy failures (WinError 112) by @anmolxlight in https://github.com/unslothai/unsloth/pull/6088
* Studio: Add Tensor-Parallel llama.cpp support by @oobabooga in https://github.com/unslothai/unsloth/pull/6040
* Studio: Add custom provider option to Connections by @Imagineer99 in https://github.com/unslothai/unsloth/pull/6112
* Studio: model selector and settings polish by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/6240
* Studio: login card polish and sidebar label alignment by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/6242
* Studio: pinnable plus menu items and saved prompt pins by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/6237
* Studio: bottom update banners, smooth llama.cpp progress, re-prompt after copy by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/6233
* fix(studio/responses): forward chat_template_kwargs enable_thinking to chat request by @Anai-Guo in https://github.com/unslothai/unsloth/pull/6202
* Studio: fix WSL Strix Halo GPU on reinstall (ROCDXG drop-in + system HIP before bundle) by @danielhanchen in https://github.com/unslothai/unsloth/pull/6227
* Studio: fully rounded Hub pills and refreshed menu icons by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/6248
* Studio: use px-2.5 for Hub option menu padding by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/6249
* Studio: fix Downloaded model list disappearing and order it by last download by @danielhanchen in https://github.com/unslothai/unsloth/pull/6247
* Studio: new-chat shortcut, composer draft autosave, archive threads by @NilayYadav in https://github.com/unslothai/unsloth/pull/5771
* Studio: persist speculative decoding preference across restart and model switch by @oobabooga in https://github.com/unslothai/unsloth/pull/6169
* Studio: refine menu chevron, tick icon, and one-line plus-menu shape by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/6251
* Studio: serve DiffusionGemma with live in-place denoising and honest stats by @danielhanchen in https://github.com/unslothai/unsloth/pull/6250
* Studio: bundle Gemma 4 chat templates (E2B/E4B + larger) and auto-apply to unsloth/gemma-4-*-GGUF by @danielhanchen in https://github.com/unslothai/unsloth/pull/6245
* feat(studio): implement S3 dataset loading (completes #5951) by @ashzak in https://github.com/unslothai/unsloth/pull/6222
* Studio: cache MCP tool discovery instead of re-probing every chat send by @oobabooga in https://github.com/unslothai/unsloth/pull/5828
* Fix Studio S3 dataset panel layout by @wasimysaid in https://github.com/unslothai/unsloth/pull/6252
* Attach DiffusionGemma visual-server from the prebuilt bundle by @danielhanchen in https://github.com/unslothai/unsloth/pull/6254

## New Contributors
* @jimdawdy-hub made their first contribution in https://github.com/unslothai/unsloth/pull/5909
* @dylanschroers made their first contribution in https://github.com/unslothai/unsloth/pull/6218
* @MdHussain121 made their first contribution in https://github.com/unslothai/unsloth/pull/6092
* @Ban921 made their first contribution in https://github.com/unslothai/unsloth/pull/6187
* @IrakliXYZ made their first contribution in https://github.com/unslothai/unsloth/pull/5967
* @Agnibha007 made their first contribution in https://github.com/unslothai/unsloth/pull/5617

**Full Changelog**: https://github.com/unslothai/unsloth/compare/v0.1.451-beta...v0.1.46-beta