v3.1

KevinVandy/mantine-react-tablev3.1Jul 25, 2024by Flux9665

AI Summary

Enhances the 7000-language TTS model with stochastic prosody generation, expanded IPA support, and improved pretraining data, alongside a revamped language similarity module.

Key Highlights

  • Stochastic prosody prediction (pitch, energy, durations)
  • Added support for more IPA modifiers
  • Increased languages in pretraining
  • Overhauled language similarity prediction modules

New Features

  • Stochastic prosody generation
  • Expanded IPA modifier support
  • Enhanced pretraining data coverage
  • Language similarity visualization overhaul

Full Release Notes

## What's Changed

This release provides new checkpoints and improves some aspects of the previous release that were not included due to time constraints.  For more information on the universal TTS model for 7000 languages, please refer to the previous release v3.0

- Prosody prediction in terms of pitch, energy and durations are now stochastic and sample from a distribution instead of assuming a one-to-one mapping. 
- Added support for more IPA modifiers to cover more languages
- Added more languages into the pretraining 
- Overhauled language similarity prediction modules and visualization

**Full Changelog**: https://github.com/DigitalPhonetics/IMS-Toucan/compare/v3.0...v3.1