v1.10.0
aandrew-me/ytDownloaderv1.10.0Mar 5, 2026by FranckyB
AI Summary
Adds comprehensive VibeVoice LoRA training and inference tools, integrates Qwen3.5 LLM models, and improves system stability by changing default Python and Gradio versions.
Key Highlights
- VibeVoice LoRA training with full parameter UI and subprocess support
- VibeVoice Streaming Presets (0.5B model) featuring 7 built-in voices
- Qwen3.5 LLM models replacing older Qwen3 versions
- Training UI enhancements including auto-save, stop button, and advanced parameter controls
Breaking Changes
- Default Python version changed to 3.11 to prevent wheel issues with 3.12
- Default Gradio version changed to 6.7
New Features
- VibeVoice LoRA finetuning with dataset validation and auto audio conversion
- VibeVoice LoRA inference with advanced generation params and smart caching
- VibeVoice Streaming 0.5B with built-in voices and KV-cache support
- Qwen3.5 LLM models added to the available list
- DeepFilterNet installation for audio denoising
Full Release Notes
## Add VibeVoice training, trained model inference, and streaming voice presets ### Training - VibeVoice LoRA finetuning via subprocess (dataset validation, auto audio conversion, train_vibevoice.jsonl generation) - Added Parameter UI: batch size, learning rate, epochs, save interval, DDPM batch multiplier, diffusion/CE loss weights, voice prompt drop, gradient accumulation, warmup steps, train diffusion head toggle, EMA on/off - Separate Qwen3 and VibeVoice parameter sections. - Auto-save/restore of all training params per engine via accordion expand - Stop Training button — terminates subprocess mid-run with clean status - Auto --train_connectors when training diffusion head ### Voice Presets — VibeVoice Trained - Load and generate with trained VibeVoice LoRA models (language model LoRA + diffusion head + connectors) - LoRA loading with PEFT task_type auto-correction (CAUSAL_LM → FEATURE_EXTRACTION) - Advanced generation params: cfg_scale, num_steps, do_sample, temperature, top_k, top_p, repetition_penalty - Trained model caching with smart reload on checkpoint change - Added option to scale the Lora effect, when applied to a sample ### Voice Presets — VibeVoice Streaming (Fast Generation) - VibeVoice Streaming 0.5B with 7 built-in preset voices (Carter, Davis, Emma, Frank, Grace, Mike, Samuel) - Auto-downloads voice prompt .pt files from GitHub, caches locally - Streaming generation params: cfg_scale, ddpm_steps - Voice prompt KV-cache for fast repeated generation ### LLM — Added Qwen3.5 models to list. - As Llama.cpp is compatible with Qwen3.5, replace Qwen3 with newer 3.5 models. ### Bug Fixes - Added optional DeepFilterNet installation for audio denoising, as well as made python3.11 the new default to prevent issues with missing wheels with python 3.12 - Fixed visibility issues in Gradio, but also tweaked requirement to make Gradio 6.7 the default - Improved Progress Notification with VibeVoice.