Russian scientists have created a compact add-on for generative neural networks that helps AI correctly convey the flow of time in video. It controls the frame rate and dynamics of events so that fast actions are not stretched out, and slow processes do not unexpectedly accelerate. The development can be connected to an already trained neural network. It also allows for the creation of more physically accurate synthetic data for training so-called physical AI.

Previously, the speed of what was happening in the frame was largely determined by the words in the prompt. However, directly controlling the real dynamics of the video was difficult. The new system allows the model to take into account how events unfold over time and set the right rhythm for the scene, rather than simply reproducing pre-learned patterns.
The problem is especially noticeable in modern video generation systems. They do not always correctly determine the speed of events. Therefore, fast movement can look unnaturally stretched, while slow action, on the contrary, can freeze or suddenly accelerate.
According to the scientists, this is due to the architectural features of such models. During training, it is often more important how good an individual frame looks than the physical accuracy of what is happening.
Russian mathematicians solved this problem using a compact module consisting of two independent blocks. One is responsible for the frame rate, and the second controls the dynamics of what is happening at each moment of the scene.
To train the system, the researchers prepared 40,000 video clips with different frame rates and varying dynamics. The set included both scenes with active human actions and slow natural processes.
The add-on can work in two ways. In the first case, it is connected to an already trained model to make movements in the created video smoother, without changing the overall nature of the generation.
In the second option, the module is fine-tuned together with the generative neural network. Then the system gains the ability to change not only the smoothness of movement but also the physics of what is happening in the scene.
The development was tested on the Kandinsky 5.0 Video Lite and Wan2.2 models. According to the test results, the naturalness of movements increased by 29%, visual quality by 8%, and the correspondence of the video to the text prompt by 19%. The project's source code was published by the researchers in open access on popular AI platforms.
Комментарии