v2.5

DigitalPhonetics/IMS-Toucanv2.5Apr 10, 2023by Flux9665

AI Summary

Introduces ToucanTTS, a new architecture designed for multilingual and low-resource research, accompanied by stable pretrained models and a new BigVGAN vocoder option.

Key Highlights

  • New ToucanTTS architecture.
  • Stable training with few datapoints.
  • New BigVGAN vocoder option (faster on GPU).

New Features

  • ToucanTTS architecture
  • BigVGAN vocoder support
  • Pretrained models
  • Improved training stability

Full Release Notes

We pack a bunch of designs into a new architecture, which will be the basis for our multilingual and low-resource research going forward. We call it ToucanTTS and as usual, provide pretrained models. The synthesis quality is very good and the training is very stable and requires few datapoints for training from scratch and even fewer for finetuning. It is hard to quantify these stats, so it's probably best to try it out yourself.

We also offer the option to use a BigVGAN vocoder, which sounds very nice, but is a bit slow on CPU. On GPU it is definitely recommended to use the new vocoder.