May-2025
unslothai/unslothMay-2025May 2, 2025by danielhanchen
AI Summary
Adds comprehensive support for Qwen3 models, including the 30B MoE variant, and enables full-finetuning and 8bit finetuning.
Key Highlights
- Full Qwen3 support (14B, 30B MoE)
- Full-finetuning and 8bit finetuning support
- GGUF saving capability
- Custom `auto_model` support for wider compatibility
New Features
- Qwen3 support
- MoE support
- Full-finetuning
- GGUF saving
- auto_model support
Full Release Notes
## Qwen 3 support + bug fixes
Please update Unsloth via `pip install --upgrade --force-reinstall unsloth unsloth_zoo`
Qwen3 notebook: https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3_(14B)-Reasoning-Conversational.ipynb
**_GRPO_** with Qwen3 notebook: https://colab.research.google.com/github/unslothai/notebooks/blob/main/nb/Qwen3_(4B)-GRPO.ipynb
<a href="https://docs.unsloth.ai/basics/qwen3-how-to-run-and-fine-tune"><img src="https://github.com/user-attachments/assets/7084aa42-b481-442b-b1a8-ccd1d3ca71ec" width="600"></a>
There are also many bug fixes in this release!
The 30B MoE is also fine-tunable in Unsloth!
```python
from unsloth import FastModel
import torch
model, tokenizer = FastModel.from_pretrained(
model_name = "unsloth/Qwen3-30B-A3B",
max_seq_length = 2048, # Choose any for long context!
load_in_4bit = True, # 4 bit quantization to reduce memory
load_in_8bit = False, # [NEW!] A bit more accurate, uses 2x memory
full_finetuning = False, # [NEW!] We have full finetuning now!
# token = "hf_...", # use one if using gated models
)
```
## What's Changed
* GGUF saving by @danielhanchen in https://github.com/unslothai/unsloth/pull/2017
* Gemma 3 readme by @danielhanchen in https://github.com/unslothai/unsloth/pull/2019
* Update README.md by @danielhanchen in https://github.com/unslothai/unsloth/pull/2028
* bug fix #2008 - load_in_4bit = True + fast_inference = True by @void-mckenzie in https://github.com/unslothai/unsloth/pull/2039
* unsloth_fast_generate model is not defined fix by @KareemMusleh in https://github.com/unslothai/unsloth/pull/2051
* Ensure trust_remote_code propagates down to unsloth_compile_transformers by @CuppaXanax in https://github.com/unslothai/unsloth/pull/2075
* Show `peft_error` by @IsaacBreen in https://github.com/unslothai/unsloth/pull/2080
* Add generation prompt error message change by @KareemMusleh in https://github.com/unslothai/unsloth/pull/2046
* Many bug fixes by @danielhanchen in https://github.com/unslothai/unsloth/pull/2087
* fix: config.torch_dtype in LlamaModel_fast_forward_inference by @lurf21 in https://github.com/unslothai/unsloth/pull/2091
* Updating new FFT 8bit support by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/2110
* Bug fixes by @danielhanchen in https://github.com/unslothai/unsloth/pull/2113
* Small fix by @danielhanchen in https://github.com/unslothai/unsloth/pull/2114
* fix(utils): add missing importlib import to fix NameError by @naliazheli in https://github.com/unslothai/unsloth/pull/2134
* Add QLoRA Train and Merge16bit Test by @jeromeku in https://github.com/unslothai/unsloth/pull/2130
* Fix Transformers 4.45 by @danielhanchen in https://github.com/unslothai/unsloth/pull/2151
* Bug Fixes by @danielhanchen in https://github.com/unslothai/unsloth/pull/2197
* Issues templates by @jeromeku in https://github.com/unslothai/unsloth/pull/2242
* Fix feature_request ISSUE_TEMPLATE by @jeromeku in https://github.com/unslothai/unsloth/pull/2250
* Registry refactor by @jeromeku in https://github.com/unslothai/unsloth/pull/2255
* Update README.md by @Kimizhao in https://github.com/unslothai/unsloth/pull/2267
* Update README.md by @jackswl in https://github.com/unslothai/unsloth/pull/2119
* Update bug_report.md by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/2323
* feat: Support custom `auto_model` for wider model compatibility (Whisper, Bert,etc) & `attn_implementation` support by @Etherll in https://github.com/unslothai/unsloth/pull/2263
* fix: improved error handling when llama.cpp build fails by @Hansehart in https://github.com/unslothai/unsloth/pull/2358
* Revert "fix: improved error handling when llama.cpp build fails" by @shimmyshimmer in https://github.com/unslothai/unsloth/pull/2375
* Fix saving 4bit for VLM by @Erland366 in https://github.com/unslothai/unsloth/pull/2381
* [WIP] Initial support for Qwen3. Will udpate when the model is released by @Datta0 in https://github.com/unslothai/unsloth/pull/2211
* Fixup qwen3 by @Datta0 in https://github.com/unslothai/unsloth/pull/2423
* Fixup qwen3 qk norm by @Datta0 in https://github.com/unslothai/unsloth/pull/2427
* Qwen3 inference fixes by @Datta0 in https://github.com/unslothai/unsloth/pull/2436
* Update mapper.py to add Qwen3 base by @Etherll in https://github.com/unslothai/unsloth/pull/2439
* Qwen 3, Bug Fixes by @danielhanchen in https://github.com/unslothai/unsloth/pull/2445
## New Contributors
* @void-mckenzie made their first contribution in https://github.com/unslothai/unsloth/pull/2039
* @CuppaXanax made their first contribution in https://github.com/unslothai/unsloth/pull/2075
* @IsaacBreen made their first contribution in https://github.com/unslothai/unsloth/pull/2080
* @lurf21 made their first contribution in https://github.com/unslothai/unsloth/pull/2091
* @naliazheli made their first contribution in https://github.com/unslothai/unsloth/pull/2134
* @jeromeku made their first contribution in https://github.com/unslothai/unsloth/pull/2130
* @Kimizhao made their first contribution in https://github.com/unslothai/unsloth/pull/2267
* @jackswl made their first contribution in https://github.com/unslothai/unsloth/pull/2119
* @Etherll made their first contribution in https://github.com/unslothai/unsloth/pull/2263
* @Hansehart made their first contribution in https://github.com/unslothai/unsloth/pull/2358
**Full Changelog**: https://github.com/unslothai/unsloth/compare/2025-03...May-2025