v1.1.0
tobias-kirschstein/ggheadv1.1.0May 7, 2024by FrankLeeeee
AI Summary
This release significantly expands the project's capabilities by introducing a robust data processing pipeline and expanding model training to support diverse resolutions and durations. It also enhances the underlying architecture to handle longer text inputs and enables advanced features like video editing and image animation.
Key Highlights
- Release of an automatic data processing pipeline for dataset building
- Improved ST-DiT architecture with rope positional encoding and qk norm
- Expanded support for various resolutions, aspect ratios, and durations
- New capabilities for video editing and image animation
New Features
- Automatic data processing pipeline
- Improved ST-DiT architecture (rope, qk norm)
- Flexible resolution and duration support
- Video editing
- Image animation
Full Release Notes
📍 Open-Sora 1.1 released - 🌠 Model weights are available [here](https://github.com/hpcaitech/Open-Sora). It is trained on 0s~15s, 144p to 720p, various aspect ratios videos. See our [report 1.1](https://github.com/hpcaitech/docs/report_02.md) for more discussions. - 🔧 Data processing pipeline v1.1 is released. An automatic [processing pipeline](https://github.com/hpcaitech/Open-Sora#data-processing) from raw videos to (text, video clip) pairs is provided, including scene cutting, filtering(aesthetic, optical flow, OCR, etc.), captioning managing. With this tool, you can easily build your video dataset. - ✅ Improved ST-DiT architecture includes rope positional encoding, qk norm, longer text length, etc. - ✅ Support training with any resolution, aspect ratio, and duration (including images). - ✅ Support image and video conditioning and video editing, and thus support animating images, connecting videos, etc. Visit the [Open-Sora Gallery](https://hpcaitech.github.io/Open-Sora/) to view more samples.