The era of AI-generated video has begun
The AI startup behind Stable Diffusion is now testing its generative AI tool for video. It can produce videos from a single still image.
12punto
We have seen AI models that produce video before, but Stability AI, the AI startup behind Stable Diffusion, is entering this field and raising the quality bar significantly. The new Stable Video Diffusion model, which stands out from its peers in terms of quality and realism, allows users to create videos from a single image.
AI CAN NOW ALSO CREATE VIDEO
As reported by Donanımhaber, Stable Video Diffusion actually comes in two model forms: SVD and SVD-XT. The first, SVD, converts still images into 576×1024 videos at 14 frames. SVD-XT uses the same architecture but increases the frames to 24. Both can produce videos at between three and 30 frames per second. These video outputs appear to be either on par with or of higher quality than the outputs obtained from Meta's latest video generation model, as well as those from Google and AI startups Runway and Pika Labs.
Alongside these developments, Stable Video Diffusion is currently only available for research purposes, not for real-world or commercial applications. While Stability AI states that potential users can sign up for a waitlist for access, the tool could be used in potential applications across advertising, education, entertainment, and many other sectors.
THERE ARE SHORTCOMINGS
The examples shown in the released video appear to be of relatively high quality and match competing generative systems. However, as the company writes, there are some shortcomings: it produces relatively short videos (less than 4 seconds), lacks perfect photorealism, cannot perform camera movements other than slow pans, has no text control, cannot produce legible text, and may not be able to properly generate people and faces.
On the training side, Stability AI says the tool was trained on a dataset of millions of videos and then underwent fine-tuning on a smaller dataset of several hundred thousand to one million videos. Stability AI emphasizes that it only used videos that are publicly available for research purposes.
Just as text-to-image AI tools have rapidly evolved to reach photorealistic levels, AI that generates video will also be able to produce much more realistic content quickly. All of this comes with risks of deepfakes, copyright issues, and various forms of misuse. Therefore, it is essential that developments are made within limitations.