"Sber" taught Kandinsky to create videos with speech, music, and ambient sounds simultaneously

The model generates videos up to five seconds long, supports Full HD, and is available for free in "GigaChat"

The Russian model Kandinsky 6.0 Video has learned to create videos with accompanying sound. The neural network generates human speech, music, and ambient sounds and synchronizes them with what is happening in the frame. This new feature is available for free to GigaChat users.

Image source: ChatGPT

To create a video, simply describe the desired scene with text or upload an image and add sound preferences. The model automatically generates the video sequence and audio track, and if there is speech, it synchronizes the voice with the character's lip movements. The maximum resolution after processing reaches Full HD — 1920×1080 pixels.

Kandinsky 6.0 Video consists of separate video and audio processing streams that work together to ensure the sound matches what is happening in the frame in terms of timing and meaning. Developers have also improved the transmission of movements of people, animals, and objects to make the generated scenes look more natural.

However, the technology currently has a significant limitation: one video with sound lasts no more than five seconds. Developers explain this by the high computational load and expect to increase the generation duration over time.

In addition to the user service, developers have open-sourced Kandinsky 6.0 models under the MIT license. The family includes a lightweight Lite version with 3 billion parameters and a larger Pro version with 29 billion parameters, and a separate model enhances the resolution of the finished video to Full HD.

Read more on the topic: