v3.1.2

hexgrad/kokorov3.1.2Oct 7, 2024by Flux9665

AI Summary

Introduces a new GUI for precise control over utterance pitch and duration, enabling iterative refinement of synthesized speech across 7000+ languages.

Key Highlights

  • New GUI allows precise manipulation of pitch and duration values.
  • Voice exchange capability while retaining intonation and duration settings.
  • Support for 7000+ languages via pretrained models.
  • Pretrained models available on Hugging Face.

New Features

  • GUI for precise control and real-time modification.
  • Voice exchange functionality.
  • Updated requirements and run_advanced_GUI_demo.py script.

Full Release Notes

This release includes a new GUI that allows you to control exactly how an utterance sounds.

You can generate a bunch of different realizations until you get one that you like. Then you can modify it further by dragging around the pitch values and the durations of individual phones. You can also exchange the voice for a different one while keeping your changes to the intonation and duration exactly as they are. And of course you can do so in over 7000 languages.

Just update the new requirements and run the `run_advanced_GUI_demo.py` script. By default it will load the pretrained models from Hugging Face🤗, but you can also specify our own. 

![image](https://github.com/user-attachments/assets/25872d43-7297-4a78-8815-5b9f267f9552)