v1.1.0
leptos-rs/leptosv1.1.0May 7, 2024by FrankLeeeee
AI Summary
V1.1.0 introduces a robust data processing pipeline, improved ST-DiT architecture with specific enhancements, and expanded support for flexible training and video editing capabilities.
Key Highlights
- Data processing pipeline v1.1 released
- Improved ST-DiT architecture (rope, qk norm, longer text)
- Support training with any resolution, aspect ratio, and duration
- Support image and video conditioning and video editing
- Model weights available for 0s~15s videos
New Features
- Automatic processing pipeline from raw videos to pairs
- Rope positional encoding and qk norm in ST-DiT
- Flexible training for various resolutions and durations
- Image and video conditioning
- Video editing and animation capabilities
Full Release Notes
📍 Open-Sora 1.1 released - 🌠 Model weights are available [here](https://github.com/hpcaitech/Open-Sora). It is trained on 0s~15s, 144p to 720p, various aspect ratios videos. See our [report 1.1](https://github.com/hpcaitech/docs/report_02.md) for more discussions. - 🔧 Data processing pipeline v1.1 is released. An automatic [processing pipeline](https://github.com/hpcaitech/Open-Sora#data-processing) from raw videos to (text, video clip) pairs is provided, including scene cutting, filtering(aesthetic, optical flow, OCR, etc.), captioning managing. With this tool, you can easily build your video dataset. - ✅ Improved ST-DiT architecture includes rope positional encoding, qk norm, longer text length, etc. - ✅ Support training with any resolution, aspect ratio, and duration (including images). - ✅ Support image and video conditioning and video editing, and thus support animating images, connecting videos, etc. Visit the [Open-Sora Gallery](https://hpcaitech.github.io/Open-Sora/) to view more samples.