v1.10.0
yukangcao/DreamAvatarv1.10.0Mar 5, 2026by FranckyB
AI Summary
This update introduces VibeVoice LoRA support, streaming voice presets, and support for Qwen3.5 LLM models, along with various UI and bug fixes.
Key Highlights
- Added VibeVoice LoRA training pipeline with extensive parameter controls
- Introduced 7 built-in streaming voice presets (0.5B model)
- Replaced Qwen3 with Qwen3.5 models for LLM support
- Integrated DeepFilterNet for audio denoising
- Made Python 3.11 the new default to prevent wheel compatibility issues
New Features
- VibeVoice LoRA training and inference
- VibeVoice Streaming voice presets
- Qwen3.5 LLM support
- DeepFilterNet audio denoising
- Gradio visibility fixes
Full Release Notes
## Add VibeVoice training, trained model inference, and streaming voice presets ### Training - VibeVoice LoRA finetuning via subprocess (dataset validation, auto audio conversion, train_vibevoice.jsonl generation) - Added Parameter UI: batch size, learning rate, epochs, save interval, DDPM batch multiplier, diffusion/CE loss weights, voice prompt drop, gradient accumulation, warmup steps, train diffusion head toggle, EMA on/off - Separate Qwen3 and VibeVoice parameter sections. - Auto-save/restore of all training params per engine via accordion expand - Stop Training button — terminates subprocess mid-run with clean status - Auto --train_connectors when training diffusion head ### Voice Presets — VibeVoice Trained - Load and generate with trained VibeVoice LoRA models (language model LoRA + diffusion head + connectors) - LoRA loading with PEFT task_type auto-correction (CAUSAL_LM → FEATURE_EXTRACTION) - Advanced generation params: cfg_scale, num_steps, do_sample, temperature, top_k, top_p, repetition_penalty - Trained model caching with smart reload on checkpoint change - Added option to scale the Lora effect, when applied to a sample ### Voice Presets — VibeVoice Streaming (Fast Generation) - VibeVoice Streaming 0.5B with 7 built-in preset voices (Carter, Davis, Emma, Frank, Grace, Mike, Samuel) - Auto-downloads voice prompt .pt files from GitHub, caches locally - Streaming generation params: cfg_scale, ddpm_steps - Voice prompt KV-cache for fast repeated generation ### LLM — Added Qwen3.5 models to list. - As Llama.cpp is compatible with Qwen3.5, replace Qwen3 with newer 3.5 models. ### Bug Fixes - Added optional DeepFilterNet installation for audio denoising, as well as made python3.11 the new default to prevent issues with missing wheels with python 3.12 - Fixed visibility issues in Gradio, but also tweaked requirement to make Gradio 6.7 the default - Improved Progress Notification with VibeVoice.