v0.1.0

FlashLabs-AI-Corp/FlashLabs-Chromav0.1.0Feb 23, 2024by yuantuo666

AI Summary

This initial stable release establishes the core architecture for Amphion, featuring implementations of VALLE, VITS-SVC, and NaturalSpeech2, along with essential workflows for training and evaluation of these audio generation models.

Key Highlights

  • Core implementation of VALLE (Text-to-Speech)
  • VITS-SVC implementation for voice conversion
  • NaturalSpeech2 integration
  • Pre-trained models availability
  • Code formatting and validation workflows

New Features

  • VALLE
  • VITS-SVC
  • NaturalSpeech2
  • Whisper feature extraction
  • Dynamic batch size

Full Release Notes

## What's Changed
* update README.md by @lmxue in https://github.com/open-mmlab/Amphion/pull/1
* Amphion Alpha Release by @RMSnow in https://github.com/open-mmlab/Amphion/pull/2


**Full Changelog**: https://github.com/open-mmlab/Amphion/commits/v0.1.0-alpha

## What's Changed
* add core code of valle by @lmxue in https://github.com/open-mmlab/Amphion/pull/4
* Refactor G2P module and related process by @lmxue in https://github.com/open-mmlab/Amphion/pull/5
* Debug VITS for multi-speaker training by @lmxue in https://github.com/open-mmlab/Amphion/pull/6
* Fix bugs for extracting whisper features by @RMSnow in https://github.com/open-mmlab/Amphion/pull/7
* Resume for the SVC's vocalist pretrained ckpt by @RMSnow in https://github.com/open-mmlab/Amphion/pull/8
* Fix tts inference bugs by @lmxue in https://github.com/open-mmlab/Amphion/pull/13
* Fix bugs for FS2 inference on updated phone_extractor by @ChenX17 in https://github.com/open-mmlab/Amphion/pull/12
* Add dynamic batch size for valle by @HeCheng0625 in https://github.com/open-mmlab/Amphion/pull/17
* Revert "Add dynamic batch size for valle" by @lmxue in https://github.com/open-mmlab/Amphion/pull/18
* Add VitsSVC implementation by @viewfinder-annn in https://github.com/open-mmlab/Amphion/pull/14
* Better training recipes for evaluation and vocoder by @VocodexElysium in https://github.com/open-mmlab/Amphion/pull/20
* Add dynamic batch size for VALLE by @HeCheng0625 in https://github.com/open-mmlab/Amphion/pull/19
* Update pretrained tts models by @lmxue in https://github.com/open-mmlab/Amphion/pull/29
* update acoustic feature extractor of TTS by @lmxue in https://github.com/open-mmlab/Amphion/pull/28
* Improve the G2P LexiconModule of TTS by @treya-lin in https://github.com/open-mmlab/Amphion/pull/30
* Add workflow to check code format with black formatter by @BakerBunker in https://github.com/open-mmlab/Amphion/pull/27
* black format on changed files to make workflow work by @lmxue in https://github.com/open-mmlab/Amphion/pull/33
* split and process valid set by @lmxue in https://github.com/open-mmlab/Amphion/pull/25
* Refine Whisper and WeNet Contents Extractor by @Adorable-Qin in https://github.com/open-mmlab/Amphion/pull/32
* Fix VitsSVC model infer bug when using nsfhifigan as generator by @viewfinder-annn in https://github.com/open-mmlab/Amphion/pull/37
* Debug for issue 23 by @ChenX17 in https://github.com/open-mmlab/Amphion/pull/34
* Control the Diffusion SVC Inference Steps in Args by @RMSnow in https://github.com/open-mmlab/Amphion/pull/38
* Add NaturalSpeech2 by @HeCheng0625 in https://github.com/open-mmlab/Amphion/pull/35
* Amphion v0.1 Release by @RMSnow in https://github.com/open-mmlab/Amphion/pull/39

## New Contributors
* @treya-lin made their first contribution in https://github.com/open-mmlab/Amphion/pull/30
* @BakerBunker made their first contribution in https://github.com/open-mmlab/Amphion/pull/27

**Full Changelog**: https://github.com/open-mmlab/Amphion/compare/v0.1.0-alpha...v0.1.0