v0.1.49-beta
chaitin/SafeLinev0.1.49-betaJul 15, 2026by danielhanchen
AI Summary
This release focuses on deep personalization, offering three color palettes, custom fonts, and seven new languages, alongside safer agents and a new Vulkan backend for Intel GPUs. It also expands the model library with the new 975B parameter Inkling model and a broader Dynamic NVFP4 collection.
Key Highlights
- Deep personalization with Standard, Classic, and Minimal palettes
- Seven new display languages including French, German, and Korean
- Safer agents with four-level tool-call permission selector
- Vulkan llama.cpp backend for Intel GPU acceleration
- Native support for the new 975B parameter Inkling model
New Features
- Custom UI color palettes and fonts
- Voice settings tab for dictation and read-aloud
- Tool-call permission selector (Ask, Approve, Off, Full access)
- Vulkan llama.cpp backend for Intel GPUs
- Inkling 975B model support with up to 1M context window
- Dynamic NVFP4 collection for various Qwen and Gemma models
Full Release Notes
Hey guys we got lots of new update for Unsloth, especially customization. Unsloth now yours to personalize: three color palettes plus custom colors and fonts, seven new display languages, and a new Voice settings tab for dictation and read-aloud. Agents get safer with a four-level tool-call permission selector (Ask, Approve for me, Off, Full access) and workspace isolation, and Intel GPUs finally get GPU-accelerated inference through new Vulkan `llama.cpp` support. [Inkling](https://unsloth.ai/docs/models/inkling), a new 975B parameter (41B active) open model with up to a 1M context window. Licensed under Apache 2.0, Inkling accepts text, images, and audio and generates text. Unsloth Studio natively supports it! ## Personalize Your Studio Appearance is now fully customizable, not just light/dark (#7077): - **Three color palettes**, each with its own light and dark scheme: **Standard** (the Unsloth green), **Classic** (a neutral enterprise grey/black/white with a sparing blue accent), and **Minimal** (strictly monochrome). - **Custom accent, background, and foreground colors** through an in-app color picker (hue/saturation, hex, eyedropper), and **custom UI, heading, chat, and code fonts** with font import (woff2/woff/ttf/otf) and device/folder font detection. - Font-size, contrast, pointer-cursor, reduced-motion, and font-smoothing controls, plus a restyle of the core controls (inputs, dropdowns, buttons) and settings search. Every setting is per-mode (light and dark keep their own values) and syncs across devices. ### Seven New Languages - Added French, German, Spanish, Hindi, Arabic, Russian, and Korean display languages (#7076), joining the existing Chinese (Simplified), Japanese, and Portuguese (Brazil), and made Auto-detect (from the browser language) the new default. ### Voice Settings - New Voice settings tab for dictation, dictionary, and read-aloud (#7074). (Full voice conversation mode is still in progress and not part of this release.) - STT and TTS coming very soon! ## Tool-Call Permissions & Safer Agents Finer control and safer defaults whenever a model calls a tool or an agent runs. - The Bypass-permissions toggle becomes a four-level tool-call permission selector: **Ask** (approve every call), **Approve for me** (auto-runs read-only calls, pauses only on potentially-unsafe ones), **Off**, and **Full access** (also disables the sandbox). Covers terminal, python, web search, RAG, and MCP tools (#7079). - Agent workspace isolation and Hermes one-shot resume (#7103); `unsloth start` now warns before running an agent's remote installer (#7024) and gains `--persist` to keep and reopen agent sessions (#7014). - Live tool output streams with SSE heartbeats, with better web-page extraction and surfaced interrupted turns (#7083); binary `web_search` fetches are rejected instead of decoded into replacement characters (#7130). - Removed `AGENTS.md` / `CLAUDE.md` from install artifacts and tightened instruction-file cleanup (#7096, #7097); the seeded bootstrap password auto-fill is now scoped to loopback clients (#7131). ## Dynamic NVFP4 See our Dynamic NVFP4 article for more details: https://unsloth.ai/docs/basics/nvfp4 We also shipped a broader **[NVFP4 collection](https://huggingface.co/collections/unsloth/nvfp4)** this release: * [Qwen3.6 35B](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4) and (Qwen3.6 35B [Fast](https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4-Fast)) and [27B](https://huggingface.co/unsloth/Qwen3.6-27B-NVFP4) * [Qwen3.5 122B](https://huggingface.co/unsloth/Qwen3.5-122B-A10B-NVFP4) * [Inkling](https://huggingface.co/unsloth/inkling-NVFP4) * [GLM-4.7 Flash](https://huggingface.co/unsloth/GLM-4.7-Flash-NVFP4) * Gemma 4: [31B](https://huggingface.co/unsloth/gemma-4-31B-it-NVFP4), [26B](https://huggingface.co/unsloth/gemma-4-26B-A4B-it-NVFP4), [12B](https://huggingface.co/unsloth/gemma-4-12b-it-NVFP4), [E4B](https://huggingface.co/unsloth/gemma-4-E4B-it-NVFP4), [E2B](https://huggingface.co/unsloth/gemma-4-E2B-it-NVFP4) ## Vulkan & llama.cpp - New Vulkan `llama.cpp` backend: Intel GPUs now get GPU-accelerated inference instead of falling back to CPU (they use Vulkan, not CUDA/ROCm), reusing the existing VRAM, context auto-sizing, multi-GPU selection, and layer-offload paths. AMD users can opt in with `UNSLOTH_FORCE_VULKAN=1` (#5819). - Prebuilts resolve via the release-assets CDN to avoid GitHub API rate limits (#7086); update-apply is pinned to the exact release the banner offered (#7112); fixed the permanent GGUF "update available" on no-symlink caches (#7113); model state resyncs after a `llama.cpp` update unloads it (#6998). - Stop now interrupts a `llama.cpp` generation stalled mid-stream (#7117). ## Models & Training - DeepSeek-V4: eager attention and trainable FP8 grouped experts (#7042). - Stopped silent full-sequence training: "Train on completions" used a hardcoded template table and silently disabled masking on a miss, so DeepSeek and most LFM models trained on prompts too while the toggle said otherwise. Markers are now auto-detected from the tokenizer's chat template (#7054); also fixed manual response-template markers that never matched their rendered templates (#7062). - Honor the "none" gradient-checkpointing option instead of silently forcing "unsloth" (#7128). - Restore the process-global `torch.compile` config on torch 2.12 so gradient-checkpointing backward honors it (#7019). - Keep native RoPE scaling when extending context, carrying `rope_theta` for linear scaling (#7028). - Force-terminate a stuck training stop after a grace period (#7099). - Route new architectures automatically: `lfm2_moe` (LFM2-8B-A1B) to transformers 5.3.0 (#7040), models by `CONFIG_MAPPING_NAMES` instead of hardcoded tables (#7043), and offer the latest transformers release for brand-new architectures (#7056). - Inkling (TML) support fixes for multi-GPU B200s: KV-cache-aware context buffer sizing, tool-call healing for its typed `<|content_invoke_tool_json|>` markers, a download-worker `HF_TOKEN` fallback, and numeric reasoning-effort mapping (#7153). ## FP8, Quant & Export - Fixed fast inference on compressed-tensors FP8 models (#7025), restored dropped FP8 `weight_scale_inv` tensors on load (#6978), and guarded FP8 Triton launches with tensor device context (#6888). - Single-pass GGUF export for directly convertible outtypes in `save.py` (#7090); concise NVFP4 inference errors (#7145). - Probe xformers support on `sm_120` instead of disabling it by version (#6828). ## OpenAI-Compatible API, GGUF & MCP - Hardened OpenAI-compatible GGUF streaming (#6950) and queue local GGUF requests before `llama-server` (#7047). - `/v1/messages` accepts thinking and null/unknown content blocks, fixing resumed Claude Code / Anthropic sessions that previously 400'd (#7017); the "No model loaded" 400 now hints at Model auto-switch (#7006); Studio auto-loads the last used local model, including GGUF variant (#6966). - Render image content returned by MCP tools (#7081); persistent stdio MCP sessions keep server state across tool calls (#7080). ## Recipes, RAG & Data Prep - Fixed `SyntheticDataKit.chunk_data` emitting chunks over `max_tokens` (#7073) and guarded `smart_chunk_text` against `stride >= chunk_size` (#7126). - Fixed `FastSentenceTransformer` Qwen embedding preprocessing (#6939) and switched recipes to a writable artifact path (#7044). ## Installer, Hardware & Platform - Fixed the Windows installer torch index override (#6972) and install a torchao Windows ROCm stub in the inference worker (#7000). - Added the `UNSLOTH_SKIP_AUTOSTART` installer flag (#7093); the installer now returns proper failure exit codes (#7123). - Detect Windows Intel GPUs via the registry before WMI (#7064) and expose Windows drive roots in the folder browser (#7082). ## Polish & Reliability - Fixed Studio auto-titles for reasoning models (#7098), user-message overflow for long unbroken text (#7100), and duplicate response model labels and hover (#7049). - Recover MLX VLM image prompts (#7094); prevent an auth-monitor reload loop (#7118); fix S3 tab flashing on reload (#7106) and Hub offline status (#7129). - Quiet noisy logs, log real progress, and speed up Windows/macOS dataset prep (#7087); added a startup loading banner and muted the benign bitsandbytes ROCm warning (#7085). - UI polish: compact thinking control in narrow composers (#7150), tidier permission composer layout and Hub feed icons (#7148), toast close-button positioning (#7142), and removed the edge-fades appearance setting (#7143). - Made the Cloudflare tunnel opt-in, off by default (#7046). # Updating Unsloth To update Unsloth or install a fresh Unsloth Unsloth, use the commands below: **macOS, Linux, WSL:** ``` curl -fsSL https://unsloth.ai/install.sh | sh ``` **Windows:** ``` irm https://unsloth.ai/install.ps1 | iex ``` ## What's Changed * Bump install.sh / install.ps1 pin to unsloth>=2026.7.1 by @danielhanchen in https://github.com/unslothai/unsloth/pull/6943 * Sort chat recents by last activity by @NilayYadav in https://github.com/unslothai/unsloth/pull/6844 * Studio: render \[ \] and \( \) LaTeX delimiters in chat by @oobabooga in https://github.com/unslothai/unsloth/pull/6914 * fix: match qwen3-thinking double-newline in train_on_responses_only response pattern by @InfoSage05 in https://github.com/unslothai/unsloth/pull/6926 * Studio: stream reasoning tokens in the tool-loop generator (fixes DeepSeek thinking not streaming with a pill on) by @oobabooga in https://github.com/unslothai/unsloth/pull/6947 * Create ossf.yml by @danielhanchen in https://github.com/unslothai/unsloth/pull/6952 * Speed up Studio startup path by @wasimysaid in https://github.com/unslothai/unsloth/pull/6899 * Polish assistant message actions menu by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/6962 * Move New badge to System settings tab by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/6963 * Fix opencode Unsloth provider selection by @Imagineer99 in https://github.com/unslothai/unsloth/pull/6906 * Fix Hermes install hint on Windows by @Imagineer99 in https://github.com/unslothai/unsloth/pull/6903 * Studio: heal DiffusionGemma tool calls into structured tool_calls by @oobabooga in https://github.com/unslothai/unsloth/pull/6851 * Fix case-variant model matching and GGUF cache reuse in unsloth start by @Imagineer99 in https://github.com/unslothai/unsloth/pull/6900 * Studio: show Hugging Face address on hover for Hub and online model rows (#6382) by @danielhanchen in https://github.com/unslothai/unsloth/pull/6928 * Studio: fix currency and indentation edge cases in LaTeX rendering by @danielhanchen in https://github.com/unslothai/unsloth/pull/6957 * Add MLX backend support for CLI unsloth train by @Lyxot in https://github.com/unslothai/unsloth/pull/6709 * feat(cli): support MLX distributed inference by @Lyxot in https://github.com/unslothai/unsloth/pull/6845 * Route MLX trainer callbacks through UnslothTrainer by @Lyxot in https://github.com/unslothai/unsloth/pull/6929 * (GRPO) Fix PEFT replacement for TRL >= 1.7.0, add missing compute_aux_loss for TRL >= 1.7.0 by @marcandrelarochelle in https://github.com/unslothai/unsloth/pull/6904 * version-compat CI: fake CPU training runs for SFT/GRPO/DPO by @danielhanchen in https://github.com/unslothai/unsloth/pull/6965 * Fix OpenClaw start default to local TUI by @Imagineer99 in https://github.com/unslothai/unsloth/pull/6937 * feat: detect installed coding agent CLIs in Studio settings by @ErenAta16 in https://github.com/unslothai/unsloth/pull/6909 * Studio: don't pin transformers before the training worker activates the 5.x sidecar by @danielhanchen in https://github.com/unslothai/unsloth/pull/6968 * Studio: source CPU llama.cpp prebuilts from unslothai/llama.cpp by @oobabooga in https://github.com/unslothai/unsloth/pull/6311 * fix(studio/hub): apply repo_id length limit per segment, not whole string (#6946) by @Anai-Guo in https://github.com/unslothai/unsloth/pull/6953 * MoE LoRA: auto-target per-expert Linear experts (gpt-oss 4bit) instead of leaving them frozen by @danielhanchen in https://github.com/unslothai/unsloth/pull/6936 * Studio: fix flash-attn and torchao install on Blackwell (sm_100+) GPUs (Closes #6961) by @ThomasEricB in https://github.com/unslothai/unsloth/pull/6970 * Fix Backend CI: add has_blackwell_gpu to the mlx worker test stub by @danielhanchen in https://github.com/unslothai/unsloth/pull/6980 * Studio: allow CPU-only DiffusionGemma by granting the diffusion runner the CPU device by @danielhanchen in https://github.com/unslothai/unsloth/pull/6979 * Bump install.sh / install.ps1 pin to unsloth>=2026.7.2 by @danielhanchen in https://github.com/unslothai/unsloth/pull/6981 * Studio: render thinking blocks for safetensors inference with prefilled <think> templates by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/6816 * Remove API menu new badge by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/6983 * Fix BAD_MAPPINGS not redirecting the -unsloth-bnb-4bit dynamic quants by @vineethsaivs in https://github.com/unslothai/unsloth/pull/6949 * Fix duplicate gemma-2b-bnb-4bit key routing the base model to the instruct model by @vineethsaivs in https://github.com/unslothai/unsloth/pull/6891 * Fix to_sharegpt optional block rendering "None" for missing extra columns by @vineethsaivs in https://github.com/unslothai/unsloth/pull/6827 * Guard FP8 Triton launches with tensor device context by @ramisworld in https://github.com/unslothai/unsloth/pull/6888 * Fix per-block ID collisions and add block cleanup for unstructured uploads by @NilayYadav in https://github.com/unslothai/unsloth/pull/6944 * Stabilize floating monitor drag by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/6984 * Retry the Studio UI shutdown re-login on transient goto timeout by @danielhanchen in https://github.com/unslothai/unsloth/pull/7027 * Fix FastSentenceTransformer Qwen embedding preprocessing by @Etherll in https://github.com/unslothai/unsloth/pull/6939 * unsloth start: warn before running an agent's remote installer by @danielhanchen in https://github.com/unslothai/unsloth/pull/7024 * Restore process-global torch.compile config on torch 2.12 so gradient checkpointing backward honors it by @danielhanchen in https://github.com/unslothai/unsloth/pull/7019 * Silence torch._check_is_size FutureWarning and shim it if torch removes it by @danielhanchen in https://github.com/unslothai/unsloth/pull/7023 * unsloth start: add --persist to keep and reopen agent sessions by @danielhanchen in https://github.com/unslothai/unsloth/pull/7014 * Studio /v1/messages: accept thinking and unknown content blocks by @danielhanchen in https://github.com/unslothai/unsloth/pull/7017 * Studio: add Vulkan llama.cpp support by @oobabooga in https://github.com/unslothai/unsloth/pull/5819 * Fix Windows installer torch index override by @alkinun in https://github.com/unslothai/unsloth/pull/6972 * Fix fast inference crash on compressed-tensors FP8 models by @danielhanchen in https://github.com/unslothai/unsloth/pull/7025 * Keep native RoPE scaling when extending context; carry rope_theta for linear by @danielhanchen in https://github.com/unslothai/unsloth/pull/7028 * scripts: refresh scan_packages allowlist baseline by @danielhanchen in https://github.com/unslothai/unsloth/pull/7032 * Studio: remove dead direct_linux_release_plan path by @danielhanchen in https://github.com/unslothai/unsloth/pull/7030 * Restore dropped FP8 weight_scale_inv tensors on load by @danielhanchen in https://github.com/unslothai/unsloth/pull/6978 * Studio: resolve the repo-root MTP drafter after the MTP/ GGUF rename by @danielhanchen in https://github.com/unslothai/unsloth/pull/7031 * Studio: harden OpenAI-compatible GGUF streaming by @Apoze in https://github.com/unslothai/unsloth/pull/6950 * Studio: auto-load last used local model by @alkinun in https://github.com/unslothai/unsloth/pull/6966 * Clarify in README that -H 0.0.0.0 starts a public Cloudflare tunnel by @oobabooga in https://github.com/unslothai/unsloth/pull/7007 * CI: retry transient HTTP timeouts in Studio smoke probes by @danielhanchen in https://github.com/unslothai/unsloth/pull/7052 * fix: guard remove_special_tokens against tokenizers without a BOS token by @vineethsaivs in https://github.com/unslothai/unsloth/pull/7048 * Studio: queue local GGUF OpenAI-compatible requests before llama-server by @Apoze in https://github.com/unslothai/unsloth/pull/7047 * Studio: hint at Model auto-switch in the OpenAI "No model loaded" 400 by @oobabooga in https://github.com/unslothai/unsloth/pull/7006 * Studio: detect Windows Intel GPUs via the registry before WMI by @oobabooga in https://github.com/unslothai/unsloth/pull/7064 * Studio: route lfm2_moe (LFM2-8B-A1B) to transformers 5.3.0 by @danielhanchen in https://github.com/unslothai/unsloth/pull/7040 * Studio: route models by CONFIG_MAPPING_NAMES instead of hardcoded tables by @danielhanchen in https://github.com/unslothai/unsloth/pull/7043 * Studio: auto-detect completion masking markers, stop silent full-sequence training by @danielhanchen in https://github.com/unslothai/unsloth/pull/7054 * scripts: refresh scan_packages allowlist baseline by @danielhanchen in https://github.com/unslothai/unsloth/pull/7078 * DeepSeek-V4: eager attention and trainable FP8 grouped experts by @danielhanchen in https://github.com/unslothai/unsloth/pull/7042 * Studio: fix the manual response-template markers that never match their rendered templates by @danielhanchen in https://github.com/unslothai/unsloth/pull/7062 * Fix SyntheticDataKit.chunk_data emitting chunks over max_tokens by @winklemad in https://github.com/unslothai/unsloth/pull/7073 * Studio: startup loading banner and mute the benign bitsandbytes ROCm warning by @danielhanchen in https://github.com/unslothai/unsloth/pull/7085 * Studio: add UNSLOTH_SKIP_AUTOSTART installer flag by @danielhanchen in https://github.com/unslothai/unsloth/pull/7093 * Studio: remove AGENTS.md and CLAUDE.md from install artifacts by @danielhanchen in https://github.com/unslothai/unsloth/pull/7096 * Tighten Studio instruction-file cleanup boundaries by @danielhanchen in https://github.com/unslothai/unsloth/pull/7097 * Fix Studio user-message overflow for long unbroken text by @Lyxot in https://github.com/unslothai/unsloth/pull/7100 * fix(studio): use writable recipe artifact path by @Lyxot in https://github.com/unslothai/unsloth/pull/7044 * Fix Studio auto-titles for reasoning models by @Lyxot in https://github.com/unslothai/unsloth/pull/7098 * Studio: resync model state after a llama.cpp update unloads it by @oobabooga in https://github.com/unslothai/unsloth/pull/6998 * Studio: Fix torch_dtype deprecation warning on startup and ASR load by @oobabooga in https://github.com/unslothai/unsloth/pull/6999 * Studio: Show Run button for downloaded non-GGUF models in the Model Hub by @oobabooga in https://github.com/unslothai/unsloth/pull/7001 * Studio: install torchao Windows ROCm stub in the inference worker by @oobabooga in https://github.com/unslothai/unsloth/pull/7000 * Studio: Add rename to project chat rows by @oobabooga in https://github.com/unslothai/unsloth/pull/7005 * Probe xformers support on sm_120 instead of disabling it by version by @oobabooga in https://github.com/unslothai/unsloth/pull/6828 * Studio: render image content returned by MCP tools by @NilayYadav in https://github.com/unslothai/unsloth/pull/7081 * Studio: add French, German, Spanish, Hindi, Arabic, Russian and Korean display languages by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/7076 * scan_packages: baseline sentencepiece dup2 finding after upstream reindent by @danielhanchen in https://github.com/unslothai/unsloth/pull/7120 * Studio: persistent stdio MCP sessions so server state survives across tool calls by @NilayYadav in https://github.com/unslothai/unsloth/pull/7080 * Studio: resolve llama.cpp prebuilts via the release-assets CDN to avoid GitHub API rate limits by @danielhanchen in https://github.com/unslothai/unsloth/pull/7086 * fix(studio): prevent auth monitor reload loop by @Lyxot in https://github.com/unslothai/unsloth/pull/7118 * Studio: make Stop interrupt a llama.cpp generation stalled mid-stream by @oobabooga in https://github.com/unslothai/unsloth/pull/7117 * Unsloth: appearance palettes, customization options, and control restyle by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/7077 * Studio: pin llama.cpp update apply to the release the banner offered by @oobabooga in https://github.com/unslothai/unsloth/pull/7112 * Studio: fix the permanent GGUF "update available" on no-symlink caches by @gaurav0107 in https://github.com/unslothai/unsloth/pull/7113 * Fix S3 tab flashing on reload by @NilayYadav in https://github.com/unslothai/unsloth/pull/7106 * fix: correct DeepScaleR-1.5B 16bit mapping to its own repo by @anxkhn in https://github.com/unslothai/unsloth/pull/7088 * fix(dataprep): guard smart_chunk_text against stride >= chunk_size by @anxkhn in https://github.com/unslothai/unsloth/pull/7126 * Fix Hub offline status by @NilayYadav in https://github.com/unslothai/unsloth/pull/7129 * fix: name unsloth_vllm_standby parameter in vLLM standby error by @anxkhn in https://github.com/unslothai/unsloth/pull/7089 * Studio: expose Windows drive roots in the folder browser by @gaurav0107 in https://github.com/unslothai/unsloth/pull/7082 * Studio CI: make tool-calling SSE probes resilient to transport stalls by @danielhanchen in https://github.com/unslothai/unsloth/pull/7137 * Studio: force-terminate a stuck training stop after a grace period by @danielhanchen in https://github.com/unslothai/unsloth/pull/7099 * fix(studio): recover MLX VLM image prompts by @Lyxot in https://github.com/unslothai/unsloth/pull/7094 * fix(install): return failure exit codes by @ShiroKSH in https://github.com/unslothai/unsloth/pull/7123 * Fix bare except clauses and remove duplicate MAX_FUSED_SIZE definition by @lxcxjxhx in https://github.com/unslothai/unsloth/pull/7138 * Studio: scope the seeded bootstrap password auto-fill to loopback clients by @oobabooga in https://github.com/unslothai/unsloth/pull/7131 * fix: single-pass GGUF export for directly convertible outtypes in save.py by @dylanschroers in https://github.com/unslothai/unsloth/pull/7090 * Studio: offer the latest transformers release for brand-new architectures by @danielhanchen in https://github.com/unslothai/unsloth/pull/7056 * Fix agent workspace isolation and Hermes one-shot resume by @wasimysaid in https://github.com/unslothai/unsloth/pull/7103 * Studio: permission levels for chat tool calls (Ask, Approve for me, Off, Full access) by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/7079 * Studio: make the Cloudflare tunnel opt-in (off by default) by @LeoBorcherding in https://github.com/unslothai/unsloth/pull/7046 * Studio: honor the 'none' gradient checkpointing option in training by @oobabooga in https://github.com/unslothai/unsloth/pull/7128 * Studio: remove the edge fades appearance setting by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/7143 * Studio: quiet noisy logs, log real progress, and speed up Windows/macOS dataset prep by @danielhanchen in https://github.com/unslothai/unsloth/pull/7087 * Show concise NVFP4 inference errors by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/7145 * Fix Studio toast close-button positioning by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/7142 * Studio: add Voice settings tab (dictation, dictionary, read aloud) by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/7074 * Studio: stream live tool output with SSE heartbeats, fix web page extraction, and surface interrupted turns by @danielhanchen in https://github.com/unslothai/unsloth/pull/7083 * Studio: fix permission composer layout and Hub feed icons by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/7148 * Compact thinking control in narrow composers by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/7150 * Studio: reject binary web_search fetches instead of decoding them into replacement chars by @oobabooga in https://github.com/unslothai/unsloth/pull/7130 * Studio: fix duplicate response model labels and hover by @Etherll in https://github.com/unslothai/unsloth/pull/7049 ## New Contributors * @ErenAta16 made their first contribution in https://github.com/unslothai/unsloth/pull/6909 * @ThomasEricB made their first contribution in https://github.com/unslothai/unsloth/pull/6970 * @Apoze made their first contribution in https://github.com/unslothai/unsloth/pull/6950 * @winklemad made their first contribution in https://github.com/unslothai/unsloth/pull/7073 * @gaurav0107 made their first contribution in https://github.com/unslothai/unsloth/pull/7113 * @ShiroKSH made their first contribution in https://github.com/unslothai/unsloth/pull/7123 * @lxcxjxhx made their first contribution in https://github.com/unslothai/unsloth/pull/7138 **Full Changelog**: https://github.com/unslothai/unsloth/compare/v0.1.48-beta...v0.1.49-beta