0.2.1

hpcaitech/Open-Sora0.2.1Nov 30, 2024by edwko

AI Summary

This release introduces support for the ExLlamaV2 inference engine and integrates Whisper to automatically transcribe audio for speaker generation when no transcript is provided.

Key Highlights

  • Added support for ExLlamaV2 inference engine (PR #37).
  • Integrated Whisper-based transcription to generate speaker text automatically.
  • Updated `create_speaker` method to accept `whisper_model` and `whisper_device` parameters.

New Features

  • ExLlamaV2 support
  • Whisper integration for automatic speaker transcription
  • Configurable Whisper model and device selection

Full Release Notes

### Release Notes v0.2.1

#### New Features and Improvements:

1. **Support for ExLlamaV2**  
   - Integrated support for ExLlamaV2
   - Pull request: #37

2. **Whisper Integration for Speaker Generation**  
   - Added Whisper-based transcription for generating speakers when no transcript is provided.  
   - Suggested in: #28 
   - Now, if `transcript` is set to `None`, the text will be automatically transcribed using Whisper.  

   ```python
   def create_speaker(
       self, 
       audio_path: str, 
       transcript: str = None, 
       whisper_model: str = "turbo",
       whisper_device = None
   )
   ```