
Hugging Face Blog
· 1 min read
State of open video generation models in Diffusers
OpenAI’s Sora demo marked a striking advance in AI-generated video last year and gave us a glimpse of the potential capabilities of video generation models. The impact was immediate and since that demo, the video generation space has become increasingly competitive with major players and startups producing their own highly capable models such as Google’s Veo2, Haliluo’s Minimax, Runway’s Gen3 Alpha, Kling, Pika, and Luma Lab’s Dream Machine.
Open-source has also had its own surge of video generation models with CogVideoX, Mochi-1, Hunyuan, Allegro, and LTX Video. Is the video community having its “Stable Diffusion moment”?
This post will provide a brief overview of the state of video generation models, where we are with respect to open video generation models, and how the Diffusers team is planning to support their adoption at scale.
Specifically, we will discuss:
- Capabilities and limitations of video generation models
- Why video generation is hard
- Open video generation models
- Video generation with Diffusers
- Inference and optimizations
- Fine-tuning
- Looking ahead
Today’s Video Generation Models and their Limitations
These are today's most popular video models for AI-generated content creation
| Provider | Model | Open/Closed | License |
|---|---|---|---|
| Meta | MovieGen | Closed (with a detailed technical report) | Proprietary |
| OpenAI | Sora | Closed | Proprietary |
| Veo 2 | Closed | Proprietary | |
| RunwayML | Gen 3 Alpha | Closed | Proprietary |
| Pika Labs | Pika 2.0 | Closed | Proprietary |
| KlingAI | Kling | Closed | Proprietary |
| Haliluo | MiniMax | Closed | Proprietary |
| THUDM | CogVideoX | Open | Custom |
| Genmo | Mochi-1 | Open | Apache 2.0 |
| RhymesAI | Allegro | Open | Apache 2.0 |
| Lightricks | LTX Video | Open | Custom |
| Tencent | Hunyuan Video | Open | Custom |
Limitations:
Why is Video Generation Hard?
There are several factors we’d like to see and control in videos:
- Adherence to Input Conditions (such as a text prompt, a starting image, etc.)
- Realism
- Aesthetics
- Motion Dynamics
- Spatio-Temporal Consistency and Coherence
- FPS
- Duration
Original source
This story was published by Hugging Face Blog. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on huggingface.co


