v3.1.2
DigitalPhonetics/IMS-Toucanv3.1.2Oct 7, 2024by Flux9665
AI Summary
This release introduces a new GUI that provides precise control over how utterances sound, allowing users to generate different realizations, manipulate pitch and duration values by dragging, and exchange voices while preserving intonation and duration changes.
Key Highlights
- New GUI for precise control over speech synthesis
- Generate multiple realizations until finding one you like
- Drag to modify pitch values and durations of individual phones
- Exchange voice while keeping intonation and duration changes
- Support for over 7000 languages
New Features
- GUI for precise control of speech synthesis
- Pitch value manipulation
- Duration manipulation of individual phones
- Voice exchange capability while preserving changes
Full Release Notes
This release includes a new GUI that allows you to control exactly how an utterance sounds. You can generate a bunch of different realizations until you get one that you like. Then you can modify it further by dragging around the pitch values and the durations of individual phones. You can also exchange the voice for a different one while keeping your changes to the intonation and duration exactly as they are. And of course you can do so in over 7000 languages. Just update the new requirements and run the `run_advanced_GUI_demo.py` script. By default it will load the pretrained models from Hugging Face🤗, but you can also specify our own. 