v1.12.5
yukangcao/DreamAvatarv1.12.5Apr 8, 2026by FranckyB
AI Summary
Integrates the Fish Speech S2 Pro (4B parameter) TTS model, offering support for 80+ languages and 15,000+ inline expression tags, along with significant compilation speed-ups.
Key Highlights
- Added Fish Speech S2 Pro as a new engine supporting 80+ languages and 15,000+ tags
- Implemented Triton/Inductor GPU kernel compilation with persistent caching for faster generations
- Added model manager methods for loading and generating voice clones
- Included Windows-specific patches for compilation compatibility
New Features
- Fish Speech S2 Pro engine integration
- GPU kernel compilation with caching
- TTS model manager API
Full Release Notes
Fish Speech S2 Pro Integration and Features: Added Fish Speech S2 Pro (4B parameter TTS model) as a new engine in the UI, supporting 80+ languages and 15,000+ inline expression tags for fine-grained speech control. The model is auto-downloaded from HuggingFace, and [tag] syntax is supported and automatically stripped for other engines. The TTS model manager (tts_manager.py) now provides get_fish_speech() and generate_voice_clone_fish_speech() methods, handling model loading, reference audio encoding, kernel cache status reporting, and batch/paragraph generation with all Fish Speech parameters exposed. Performance and Stability Enhancements: Implemented Triton/Inductor GPU kernel compilation with persistent caching in models/.cache for Fish Speech, significantly accelerating repeat generations. Added Windows-specific patches for compilation compatibility and informative first-run cache warnings.