v3.0.0
lyogavin/airllmv3.0.0Jun 30, 2026by lyogavin
AI Summary
A major overhaul that enables inference of the largest open models (up to 405B parameters) on small GPUs, introduces native FP8 support, and ensures compatibility with modern Hugging Face libraries.
Key Highlights
- Run 70B models on 4GB, 405B Llama 3.1 on 8GB, and DeepSeek-V3 on ~12GB VRAM.
- Native FP8 support for recent checkpoints like DeepSeek-V3 and Qwen3-FP8.
- Support for latest model families including Qwen3, DeepSeek-V3, and Phi-4.
- Full compatibility with current `transformers` and `accelerate` releases.
New Features
- Native FP8 support.
- Support for Qwen3, DeepSeek-V3, and Phi-4 model families.
- Reworked layer streaming for better compatibility.
Full Release Notes
## AirLLM v3.0.0 Big update: AirLLM now runs today's largest open models on tiny GPUs, with full support for the latest model families and Hugging Face versions — still no quantization, distillation, or pruning required. ### Highlights - **Run the biggest open models on a single small GPU.** Stream 70B models on 4GB, 405B Llama 3.1 on 8GB, and even **DeepSeek-V3 (671B) on ~12GB**. - **Native FP8 support.** Pre-quantized FP8 (block-FP8) checkpoints now load and run correctly — including DeepSeek-V3 and the Qwen3-FP8 family. - **Latest models supported**, including Qwen3 (dense + MoE, e.g. Qwen3-32B, Qwen3-30B-A3B, Qwen3-235B-A22B-FP8), DeepSeek-V3, Phi-4, Mixtral-8x7B, and DeepSeek-V2-Lite. - **Up to date with modern Hugging Face.** Works with current `transformers` / `accelerate` releases, so a plain `pip install airllm` just works — no manual dependency juggling. ### Improvements & fixes - Reworked layer streaming to build on the standard Transformers model path for better model compatibility and `generate()` behavior. - Runtime precision now follows each model's native dtype (e.g. bfloat16) instead of being forced to fp16, fixing garbled output on very deep models. - Fixed weight loading for layers whose tensors span multiple checkpoint shards (affected large FP8/MoE models). - More robust shard naming and attention-implementation fallback. ### Install / upgrade ```bash pip install --upgrade airllm ``` See the [README](https://github.com/lyogavin/airllm#readme) for quickstart and the full list of supported models.